Pith. sign in

Paper Citation Record · LEDGER

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

As of 15 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.09128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09128 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:05:34.799690Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66099bef-a43e-4607-8aed-e7f68deef70a · outbound

This paper cites Acta Universitatis Sapientiae, Informatica , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Acta Universitatis Sapientiae, Informatica , volume=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.528855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.540529Z digest=sha256:bd068cf8cc50f68881c967fd10ee525b7aa9f30283f816f816093f481eb4da86

Observation 3daf168c-8bbb-42d9-aeae-cfefd8631438 · outbound

This paper cites NAACL , year=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments NAACL , year=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.516410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.546631Z digest=sha256:d8276ce2a4b180dfd5823b1e383b0ab2220cb78d0b4b5fe23129231802a1dff3

Observation ce18d556-0ae4-4e64-9da2-f90757690548 · outbound

This paper cites 2026 , booktitle=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2026 , booktitle=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.502952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.550779Z digest=sha256:19e429db7c2c938edbf3bdcf19c4a71606f42192214e549e0cce147f1bae7b1c

Observation 29ab7034-16d0-42d2-8406-7e74c1caee57 · outbound

This paper cites Findings of ACL , year=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Findings of ACL , year=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.489237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.554824Z digest=sha256:09cdd1d87119613f7451cce1526265ac2a3d6cbc2099d14c86e1af3b51bf46c0

Observation 594e5ffc-4f00-47f3-bef2-f6cf10b0011c · outbound

This paper cites 2024 , booktitle=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2024 , booktitle=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.477232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.559522Z digest=sha256:91b2e41054269a755fd4a6f078e0f146c41f58b7be54c2b7513e38fa812d9f11

Observation 14d77b26-a516-41de-a5a5-469d56fcc0f1 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.465794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.563464Z digest=sha256:703a4cdb82da0d62ed307dc7a33d49abaa3de0bc32127c64bbb19fd877ee77ac

Observation d74d4f7d-84bb-45ba-b997-44c7be788b18 · outbound

This paper cites Advances in neural information processing systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in neural information processing systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.567592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.567592Z digest=sha256:4a3fbb2fbe7281c6c24492fd088facd017f0785519ec6cdb9c86d3fb8c698763

Observation 6e832812-f636-4569-b8b4-254b609b2c2f · outbound

This paper cites 2025 , month = may, note =.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , month = may, note =

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.445925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.571679Z digest=sha256:cf8090c3dc9301a99a63e6a0a24eea9bfa971244c941c930b247736b0a932dfc

Observation 89396d3a-c612-4fdf-a82c-907dfb2ad00e · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.575440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.575440Z digest=sha256:e6dacfd60910f0624d9b7a66178e7aa3001f217a0898cf8561b69ad51e808aca

Observation 52c6cd20-333f-404f-b795-6f83819e7cf6 · outbound

This paper cites Societal Alignment Frameworks Can Improve LLM Alignment.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Societal Alignment Frameworks Can Improve LLM Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.579870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.579870Z digest=sha256:38474732ded36203169ced8aad3463181d8d6919024c635280af6eb0c77882d9

Observation 3675d94c-a18b-443c-9102-a62a273b4660 · outbound

This paper cites 2005 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2005 , publisher=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.584416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.584416Z digest=sha256:9117481e341dbea5b3bdaf6c749a348429e2c19a93780391560979eef5c1fd88

Observation 8b79025d-e39c-446b-94b7-b8c4220a931c · outbound

This paper cites Concrete Problems in AI Safety.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Concrete Problems in AI Safety

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.588577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.588577Z digest=sha256:ebde10eb1120bd49e27643aa65717a168b7673772a5361c90c4e3dc8b7bcce2d

Observation d894df8e-ee2f-42df-9b5e-1ddcdea9fabb · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.592692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.592692Z digest=sha256:cf073616ec4e3826865acd2cfddcfa1f7dc91d0b6a34d6122c7279ef952e4e62

Observation e58cea66-1fc2-4454-862b-38e8fe86ffcc · outbound

This paper cites Artificial intelligence , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Artificial intelligence , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.597498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.597498Z digest=sha256:7bbc41450fbe4db91bb0315edc756ca6452df22c3ccb388845929b86c65b08cf

Observation 45bf8575-057e-4261-a378-331508fe4b8b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.601259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.601259Z digest=sha256:20f5dbd34145292b5db86aaed4a8fd7f94fff656220234ef87f21f588ae8f501

Observation da6f06b8-be26-4c03-ac42-45b7dc6c5ee9 · outbound

This paper cites The Llama 3 Herd of Models.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.605581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.605581Z digest=sha256:1297c22c46bebcd7684ab1bfe6002de513b1352b7c442726014e6e5fcf9a75d3

Observation e9625301-19d0-46f4-a001-fee5262eba8c · outbound

This paper cites 2024 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2024 , eprint=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.609259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.609259Z digest=sha256:7cde02aade54478682bf4f596f747bba72aa47f1bba29be4e37ff9611c322625

Observation b0f12945-cad7-4cf5-8ef4-1249fd64fc6a · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.398675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.613177Z digest=sha256:65a27f44d7b3dc77475c3d732b38a85ef1c442fccfe6c0040fb05f481d797420

Observation f7dac24d-6f34-45eb-a0b2-19621337f3a0 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.385597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.616839Z digest=sha256:69eb0fbdad21ec71ff1da6a3a61b0812227173e30b24becac9f8da5853e119c3

Observation 954e4e6b-ef1d-49a7-b228-086e3cca04e2 · outbound

This paper cites 2026 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2026 , eprint=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.371792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.620698Z digest=sha256:adc2811b48443a6bf7d1ae1bf10acac0df2659abbabadffb3e455378208066aa

Observation 3d011c37-78ee-4b60-bb73-3d9b00f803b9 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.624472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.624472Z digest=sha256:f80c8e9045c5e326f0dda791e0202da81dfa7992ce33f2cf7ee641393f08a998

Observation 84d06b87-c67f-491f-b547-53ee684b4f78 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in Neural Information Processing Systems , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.628227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.628227Z digest=sha256:93cef7bd0466a1959c40efeac609f708c0b87821e7e5fc00000d2076e035c968

Observation 773ede4d-d5d0-4f6a-ae28-e37296867b06 · outbound

This paper cites Proceedings of the 36th annual acm symposium on user interface software and technology , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 36th annual acm symposium on user interface software and technology , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.631953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.631953Z digest=sha256:cd4b86c003708e4bdcfd5555c11cb303f77fcb78e4c312101ae0ef9a213640aa

Observation c8645620-ea0e-4777-b1cc-ccc40dd6a6ee · outbound

This paper cites Proceedings of the National Academy of Sciences , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the National Academy of Sciences , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.635924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.635924Z digest=sha256:5dfaaeee45557b489ebd19adbf00f5ea6212bee23ce871687c209cebec4832be

Observation 2894105d-3f0e-4a29-aa5b-eddbeedf496f · outbound

This paper cites Science , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Science , volume=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.329977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.639702Z digest=sha256:6e08076d24c56db964f3e36eda629251abbd520ed0c690385701e9ca473dae24

Observation bbe5a7ab-7a84-44d4-a768-0abee4856344 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.643711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.643711Z digest=sha256:9d7237391e5c7d27749fa0e29a5febb1c3414460b21d5c35ba76273118fb6611

Observation 975706fe-9bd4-4fc4-81a7-b86d4ce42419 · outbound

This paper cites Nature Human Behaviour , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Nature Human Behaviour , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.647924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.647924Z digest=sha256:8ab2830da840e4b7d3361a308da7ecdca14f4602908d9db4de6a4d31e79e1ac8

Observation 4454e18a-3ab3-4ea1-98bf-c82f13f29c35 · outbound

This paper cites Advances in neural information processing systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in neural information processing systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.651661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.651661Z digest=sha256:9eb515382f959cccd1d6b3a8401ef534371190ba3712e288f8719710fa40dde7

Observation bdf964a1-19f4-4463-a7e5-9bb85bd572c8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in Neural Information Processing Systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.655444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.655444Z digest=sha256:9554e1c1d7ad887452bdd49b2ddb08e75c84f36548d183f4783599c6b051d7f3

Observation 77601210-e2c9-4ae2-a34a-c04fb698737a · outbound

This paper cites Advances in neural information processing systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in neural information processing systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.659162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.659162Z digest=sha256:8321aab8b4abd8544f04e2e99c4d3d474a7229585ae62192df7461e2233df601

Observation 17b30183-258b-43a1-985e-8288a569860e · outbound

This paper cites LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.662836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.662836Z digest=sha256:4824e64903000393fd086f7e0028eae59af6cc03497aee100614b183b2383fff

Observation 127bc410-80ff-4c93-a5df-4d0663cca398 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.666916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.666916Z digest=sha256:3c155bc342174e8539ec990f5c8b99ef2c257ac2d5137a3e39f97c12fd9f0928

Observation 496b2712-ea19-4b53-a284-9e8520e591f6 · outbound

This paper cites 1978 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1978 , publisher=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.281839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.670736Z digest=sha256:e8c4e96b24aff4140df3b4b11bf3ed54c56083e6c57759867c8508b5de2123f6

Observation a19a1aeb-11ce-43c4-a3f4-6b1ea2129004 · outbound

This paper cites International Conference on Machine Learning , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments International Conference on Machine Learning , pages=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.270011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.674770Z digest=sha256:f95849c9bd8336e8e99fe951b9fcedf203849fcf6777a4d85dad05e2adb54faf

Observation 8dbaa87e-8e9a-4f7d-8174-a8fb22c37e6e · outbound

This paper cites the method of paired comparisons , author=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments the method of paired comparisons , author=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.678632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.678632Z digest=sha256:3e291ad1b2ff5544800dc6c10556fabcd6b0cf13c050aa2a1f0a758321698587

Observation bfa86fc2-5cce-4729-a821-5265ed3d9a80 · outbound

This paper cites Harper's Magazine , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Harper's Magazine , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.251302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.682484Z digest=sha256:fc44571dbc3e1a0d17c70e8f065715dc11c6c97af4cbbb0074e0841f5d091541

Observation 5e5c370e-7aab-42ea-b054-d79e08c65fed · outbound

This paper cites 2011 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2011 , publisher=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.686119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.686119Z digest=sha256:74ef209c42e726a339b798b2ea343da9dc05cdc0df7a2b66187f1ad877725783

Observation e78b11c9-eff8-4e9c-a444-ab8939b2b668 · outbound

This paper cites science , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments science , volume=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.231837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.689919Z digest=sha256:37bb7c3d8f245cd7b271626668a09fa0ac96abb7bf6d1ab8042f150c0b59289b

Observation 70a90aee-3bb6-4f08-9859-8440827bb1d2 · outbound

This paper cites 1980 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1980 , publisher=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.219880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.693614Z digest=sha256:f3f0fa4038da4484533351053e063ebc5ec0c335bb540c2f7958a3c20ef54e7d

Observation 9923f44d-4655-4f10-92c0-bff98c3bd522 · outbound

This paper cites 1990 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1990 , publisher=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.697208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.697208Z digest=sha256:d0475fc946ffab7f736b4e61621509791848a16b2993ad33607500b831dee94d

Observation 72eece5e-e793-4ef5-ba53-1c448cd309d5 · outbound

This paper cites American Economic Review , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments American Economic Review , volume=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.199317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.700971Z digest=sha256:541a3b42412e8b98ba88d139c36e170a5d5f72e933ebc87ffe794eae5d756381

Observation 5e426b3f-8ced-4adf-9d53-be49d5a8976c · outbound

This paper cites Journal of Economic theory , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Journal of Economic theory , volume=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.186977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.704798Z digest=sha256:184c2ee38c65b0cf85ac6199d4d96c7ee1fe9e5c87449544b61f2550fdb30572

Observation d61c3931-0e43-4924-b939-4beba99641db · outbound

This paper cites Econometrica: Journal of the Econometric Society , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Econometrica: Journal of the Econometric Society , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.174422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.708341Z digest=sha256:c4257929e716dbf92a57ead242e0931827902bd59796f7834ebef4a04eaa02b4

Observation 3b7977d3-78c1-4106-970a-0b4540b5b125 · outbound

This paper cites , author=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments , author=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.160685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.711994Z digest=sha256:8d7992872ecccaf14ee06d8c6519073ced14fb18d4b82196d466ab5b8fdc9604

Observation 8fe176d4-dc5e-4082-8cc1-b7d4313b7eb5 · outbound

This paper cites John thinks that Mary thinks that….

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments John thinks that Mary thinks that…

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.715483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.715483Z digest=sha256:952d60da8dbfc2c4d711b56eb850d47372f1a824f7ee4dc036e5f79a74a927a4

Observation 156be659-4703-4394-bce9-251668320b51 · outbound

This paper cites Evolutionary Anthropology: Issues, News, and Reviews , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Evolutionary Anthropology: Issues, News, and Reviews , volume=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.147997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.721296Z digest=sha256:942c0d1614c7a8cd211ab91a545e8b1ba67d1420484bb9a10388279d2a1244b5

Observation 191024f9-c656-491a-a870-4531d69f5959 · outbound

This paper cites 1984 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1984 , publisher=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.136750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.725068Z digest=sha256:122d0932aafa441e0a5ead64c8a732de53b0739364cd47b9f3840e0a64a0a57c

Observation c96f0ba7-6dc2-493b-beb6-1dc1959b545d · outbound

This paper cites Cognition , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Cognition , volume=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.125061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.728658Z digest=sha256:05ca78c0a488534ef9219cc32778089f1843a6b3496d9b9bd917031b46fd05d0

Observation dbe2ea54-809a-4c1a-919d-03136eda64ba · outbound

This paper cites Cognition , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Cognition , volume=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.111809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.731838Z digest=sha256:ee2df96c3d8568f9cb0fc768d20a5a5d1142eec4437db078143fc3655ac745ee

Observation 519e20e8-bb82-4bca-b07b-8c09befafff7 · outbound

This paper cites 1988 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 1988 , publisher=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.099667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.735336Z digest=sha256:7d27f200a849089857ad69c90ac04a5e5fa210e26dede9618b10259a2265cb35

Observation 2a5a11b2-97c1-415a-9ec8-739f58cfd8bf · outbound

This paper cites an unresolved cited work.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.738866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.738866Z digest=sha256:7f12d28a60aebcf9fb9358c26986f410d16b37a110631e346e6c7a211a68f582

Observation 8efb1f2c-664e-4b59-b3d8-904f78ce0a79 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.742161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.742161Z digest=sha256:90f9281275cd9aec10df1918158217beb46476fc42f1c3ff989acef119ff7403

Observation 9b046e81-6669-4544-82ea-8f22a48688d7 · outbound

This paper cites Proceedings of the 2022 conference on empirical methods in natural language processing , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 2022 conference on empirical methods in natural language processing , pages=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.073627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.745589Z digest=sha256:281399667546f52210b20bc1b445f11138b2d8eaecc271337eb0e036a2434d74

Observation 0bdd7741-8875-4777-b65a-8942b19cddd3 · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Evaluating Large Language Models in Theory of Mind Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.748777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.748777Z digest=sha256:5db1d12d169ccd9c10a5fe1f036f9d73b46ae961164f50d47c52717b81200eba

Observation bfd5295b-c9a8-40a6-878a-bc1bde0975ae · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.752241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.752241Z digest=sha256:2577dc0d4eaf6ae4beeb7fef6b1f1b6dd564435b0bf4be4e4dd444fbda489356

Observation f800c627-c57e-4d71-9de6-51a68e340350 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.755550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.755550Z digest=sha256:4e5c79ed0232e89de61dac330305d76aea2573dd4c12d4546ca132f4de363137

Observation 6f7e96a1-d04a-46b3-8b66-362b221ffcdd · outbound

This paper cites GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.759492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.759492Z digest=sha256:1148b0ea2d117a427ba6ac1254fab5d06b5acbf00e295e9638a4640d08e0276e

Observation a9219dcf-52df-49df-b870-ed3972b466cb · outbound

This paper cites International Conference on Learning Representations , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments International Conference on Learning Representations , volume=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.061007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.764350Z digest=sha256:4c6cc149307984cf1b7c6aa30b9eada34e8f27b614211e58e13c4f70711dfee0

Observation ffbaccbe-f001-4bae-aca2-42e1d6af22eb · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.767877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.767877Z digest=sha256:29d5b809a8bf99bda4819407e8d4e9ec7eb6c226c490b996128480323be9862e

Observation 0d998186-d093-463f-b125-9c66efc035c4 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.771496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.771496Z digest=sha256:a1d1f770fb2383ef060a0abc87dd2513f262e2d40a90e757850728eb93dae043

Observation 3fda412d-e11f-4999-9d0e-50d33c78b3ed · outbound

This paper cites Handbook of Intelligence , editor=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Handbook of Intelligence , editor=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.039264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.774899Z digest=sha256:dd397d91bf0647a67f97950e09f22a6e910816bee3979e840d7f5c3854ef9bbd

Observation 81347811-66ea-4fea-af28-c0484b9fe772 · outbound

This paper cites 2025 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2025 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.778497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.778497Z digest=sha256:47c785849cf0ceb73d9af8e3584f3e00f20d8bcb0efdff8f9b0e39d40e43f92e

Observation e596d2f7-e9ad-4f6c-ac81-4781c8d7da5d · outbound

This paper cites 2023 , eprint=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2023 , eprint=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.781939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.781939Z digest=sha256:c188c010a2824bbf5108fb4a9e96b9df8df73070a49d8a76df01ae6b357c0d40

Observation 8898c9df-a433-4473-89cc-42766f86ecae · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:35.011999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.785370Z digest=sha256:9c85d5e25a81ff7f4cb6d83b78d0d35d10fd4cfdb0730dfef47bb3905b3a6e70

Observation 5c155570-937a-4238-9a6d-c70946b3ad1f · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.788809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.788809Z digest=sha256:6e1abe31e8fa2aea2159193a016d99cd71ff8a8c6fa54c3acd0436859e6e4768

Observation 2cbbefed-68ef-4c45-9eea-87278b528cb0 · outbound

This paper cites an unresolved cited work.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:34.792530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:34.792530Z digest=sha256:50152ceada0de36246f3caecec1b6ef864aada9c9f0012dede4296db6c4624b8

Observation 77bc36d6-7d6d-482c-a4b2-f9903d3755ec · outbound

This paper cites 2024 , publisher=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments 2024 , publisher=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:34.991239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.796172Z digest=sha256:1b0139172308c9bc7d1ea7d554d70f77068bd99fe47a9c8e03d11dbea0e0290f

Observation 89df69cd-4405-40d8-a20b-7102f040f3d0 · outbound

This paper cites Journal of Consumer Research , volume=.

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Journal of Consumer Research , volume=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:34.979116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T23:05:34.799690Z digest=sha256:26757a715ac9320df86ae1d2f06ded7ac4a0dd33d9a32a634237811ac9e57d02

Pith citing papers

No inbound Pith citation observations are available.