Pith. sign in

Paper Citation Record · LEDGER

lmgame-Bench: How Good are LLMs at Playing Games?

As of 19 August 2026, this Paper Citation Record lists 100 of 115 outbound references and 22 inbound Pith citation observations for arXiv:2505.15146.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15146 v2

Coverage vector

measured 100 of 115 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:27:03.187728Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.015304Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 115 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved81
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a7e70288-ebc0-4e09-82a6-f0ea6f228c64 · outbound

This paper cites OpenAI Gym.

lmgame-Bench: How Good are LLMs at Playing Games? OpenAI Gym

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:55.977387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:55.977387Z digest=sha256:abdb2fa693380798c08b35123a43c5a6f1b0e4f717701d3ad756220e945248e2

Observation 131c4169-85b2-400a-865b-adda2c05762f · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

lmgame-Bench: How Good are LLMs at Playing Games? Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.072057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.072057Z digest=sha256:a7031ff0bd0e94597d0061ef0412e32dcaea9b1cf4ad481950f8655c2ed09ced

Observation ba224b0f-9985-4a21-bb0e-6fa7100a6c5c · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

lmgame-Bench: How Good are LLMs at Playing Games? RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.140833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.140833Z digest=sha256:a86d5baa63b659bb614fb939ab07a3a839c5c0bcf74c9211af32b3ace58716ee

Observation 14a9e272-3755-4f82-8ebb-4c5d0b603de9 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

lmgame-Bench: How Good are LLMs at Playing Games? SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.255061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.255061Z digest=sha256:8573c1a41d60a4511dd4e1eeed4770cb3a23df3b7641fd42ec8a1efd61519ba2

Observation 1d9d2fe6-ad2e-4a8c-ae80-807da734ae0b · outbound

This paper cites In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., eds.: Advances in Neural Information Processing Systems.

lmgame-Bench: How Good are LLMs at Playing Games? In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., eds.: Advances in Neural Information Processing Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.363888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.363888Z digest=sha256:210af1a587bffb65848b577df61e7dba6edcec08a7c3c20cd6143455e52f27d1

Observation 6d0d7fbc-84b1-43ad-a74d-6b37e571ae55 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.453904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.453904Z digest=sha256:e8a675e16405c6ab54839bde91bd75d1ed4513aba4eae5b32c354cdf3f83c464

Observation babc8d0b-adff-4d21-9ae3-ee3d6a759645 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.552570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.552570Z digest=sha256:9a8608b670878f92cd8de5b5bbc799247a0c391c862caff232c35ddaa4a61fee

Observation 57530e12-cda5-41fe-bf7f-6574601f213e · outbound

This paper cites EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges.

lmgame-Bench: How Good are LLMs at Playing Games? EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.639836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.639836Z digest=sha256:b0834a0ae747f288321e49bd2ed7736a0b6211575df7d8ad04b065af36600c39

Observation a0ca5e66-1c30-491b-9df9-19c7522ef01d · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.720112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.720112Z digest=sha256:03988d50c611a844cf2d5c76dbe7e774669853c1314729a9a7b7dfe85338950f

Observation fc84ddec-7530-45b2-aa78-08091fdbc4a6 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

lmgame-Bench: How Good are LLMs at Playing Games? BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.816724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.816724Z digest=sha256:0be88cd8e07dc3bde4c4f7d16e4a3434b9f65d8a4fafa73ee368ca545e09a896

Observation 7e38bd61-6f22-4a25-8709-7c5227b6d909 · outbound

This paper cites GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents.

lmgame-Bench: How Good are LLMs at Playing Games? GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.894546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.894546Z digest=sha256:65538258b9f4d6171c7f308979ba29e0a41b01b0408e586bff477cda50747cb4

Observation feb50f57-f0b3-4573-880d-350a08c99f3b · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

lmgame-Bench: How Good are LLMs at Playing Games? SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.020933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.020933Z digest=sha256:be9ebf0dd1a96d9556dd362b730405cb30bb39fbe086f945bec818750d0c2246

Observation 6aa6771e-9efd-40a8-8ccd-f8cd596c8f53 · outbound

This paper cites IEEE Transactions on Games11(3) (2019) 195–202.

lmgame-Bench: How Good are LLMs at Playing Games? IEEE Transactions on Games11(3) (2019) 195–202

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.127873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.127873Z digest=sha256:662fd2d53030efce6e38dd526ffe1e0d609b91f709c997be72241101a911b247

Observation 97837f3f-1fdd-4b8d-9c47-82e019956ede · outbound

This paper cites AI Magazine22(2) (2001) 15–25.

lmgame-Bench: How Good are LLMs at Playing Games? AI Magazine22(2) (2001) 15–25

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.204676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.204676Z digest=sha256:9d7554d3556ed6168998d1fd5824bf198183a9753c40f471c8514bb1e6f1f9e9

Observation 14041c05-4acb-4b6b-bb51-360566b36a2b · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

lmgame-Bench: How Good are LLMs at Playing Games? Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.341174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.341174Z digest=sha256:80ab13ab66bb2e0d987120a71c1072c6665610f080bbeaf90456c6d2c64d9389

Observation 2607c706-b810-4f3e-a02e-76c5f1a9e56e · outbound

This paper cites Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games.

lmgame-Bench: How Good are LLMs at Playing Games? Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.413647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.413647Z digest=sha256:c77d3d2200ab066e6c726676cce45ea61a9c66fe19fbfdb8640ead96e0b83e17

Observation 42618b84-232d-4f90-b0d7-8e954d3735c2 · outbound

This paper cites Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot.

lmgame-Bench: How Good are LLMs at Playing Games? Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:27:03.886471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:26:57.507753Z digest=sha256:d71b45704f2a8763c2332ec07187a6cb697275103b271d23dac284e79798cc10

Observation 0e90353b-7af8-4f73-8b05-213dc2a49d13 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.604725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.604725Z digest=sha256:d06d164cb368df4b3d7f8403cada968cb5e8cecaa884253ca8db6711c0097cc9

Observation 3c06fd2b-4bce-456f-ae46-6b275c790620 · outbound

This paper cites OpenAI o1 System Card.

lmgame-Bench: How Good are LLMs at Playing Games? OpenAI o1 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.709482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.709482Z digest=sha256:a8566ff587c341646ed269e44b9edc22fc32b405421acfd0603516b6fc38adc2

Observation e2613515-e9ef-4538-b0cb-848ce7429f6a · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.792901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.792901Z digest=sha256:67a83a2879f38c95360dc72f5c157caa2e34f7930723b6c03a46f625ea7abaaa

Observation f282f996-6b9a-4e1f-8578-ad4a26aa3d85 · outbound

This paper cites In ICAPS.

lmgame-Bench: How Good are LLMs at Playing Games? In ICAPS

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.864760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.864760Z digest=sha256:5c0b22965bb07b0cccdaad329f0e73d1bd4c92219e056445bc5706cb5150e8bc

Observation 5ed9dfa1-d833-40fa-af0e-82ebf70154f0 · outbound

This paper cites Applied cognitive psychology31(4) (2017) 438–445.

lmgame-Bench: How Good are LLMs at Playing Games? Applied cognitive psychology31(4) (2017) 438–445

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.936132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.936132Z digest=sha256:d6ea9ae241f643296cf5b04de2bcd50d95d9ea0e59b981e94b2a329eaed2b6b7

Observation a3791a52-c555-4ec5-bb2f-a3ea2101db68 · outbound

This paper cites In International Computing and Combinatorics Conference (COCOON).

lmgame-Bench: How Good are LLMs at Playing Games? In International Computing and Combinatorics Conference (COCOON)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.046424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.046424Z digest=sha256:20450d08ca6c2801607830c8968ff6b4490213122d20f8401e9ee8583c29a42c

Observation a661cdd9-5b21-433f-b165-7d7818203605 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.163042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.163042Z digest=sha256:5c1ece0f009a542ceb5057fa682a5e75f925e246bbbd278ea194a05eef483881

Observation 458995a7-b87c-48cf-892d-495c8af4cd2b · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.220901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.220901Z digest=sha256:01b0c5c07dfc9ad5415eb37e05a0c553d133d9d167d74fadd07053a4911ac9e4

Observation 68d50081-a23e-4b11-ad68-4a9665b646da · outbound

This paper cites arXiv (2024).

lmgame-Bench: How Good are LLMs at Playing Games? arXiv (2024)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.280311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.280311Z digest=sha256:abd49c633c0500f24df128d1ec35892e5454b4e731ea9977c898023b8f5f8c98

Observation eadfffa4-9db8-4542-9f36-cc139a0416a3 · outbound

This paper cites arXiv (2025).

lmgame-Bench: How Good are LLMs at Playing Games? arXiv (2025)

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.353668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.353668Z digest=sha256:b3d5d147480da01a1298af6df7c335f7af89638f319eeb7f5f84696c4621e907

Observation b9973ff4-17fc-4c65-a76e-ebac676753b3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

lmgame-Bench: How Good are LLMs at Playing Games? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.427826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.427826Z digest=sha256:e212ca7ec631cbdce24932d854ca5b42392beaaf82d391a8f59a561a2b97fdd3

Observation 24d32840-27ff-4a3a-95c4-4e9206bb84f4 · outbound

This paper cites Computational Geometry 13(4) (1999) 215–228.

lmgame-Bench: How Good are LLMs at Playing Games? Computational Geometry 13(4) (1999) 215–228

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.487599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.487599Z digest=sha256:d5b460d9ec8d6997fe360aec797a9874ee5fa6add09d463b3e15d31efd54f666

Observation 315386b0-a926-4c90-b7d8-da2b71f4b8b2 · outbound

This paper cites A Simple Family of Analytical Trumpet Slices of the Schwarzschild Spacetime.

lmgame-Bench: How Good are LLMs at Playing Games? A Simple Family of Analytical Trumpet Slices of the Schwarzschild Spacetime

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:27:03.824241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:26:58.543471Z digest=sha256:fa71f1beb64591e05f57c199c6967761499fd6ace452abd978baa4d9e47644cc

Observation df103d95-f26f-4379-a55b-5e402af91911 · outbound

This paper cites Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models.

lmgame-Bench: How Good are LLMs at Playing Games? Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.602648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.602648Z digest=sha256:ec05848932e8c2cb5f42f0c8ab86fad140921a8d6eb8db607caa6eb1fc0b45ea

Observation b6f11591-6ffc-4ad2-a2a3-a1758280f791 · outbound

This paper cites The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks.

lmgame-Bench: How Good are LLMs at Playing Games? The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.673836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.673836Z digest=sha256:5f1ff1e0b37af2a6556cdd8b9539af8c77e4c10ae12c56a4c32dc40e4873ecc3

Observation 779eca80-188b-4c50-852e-178d48195594 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.734689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.734689Z digest=sha256:330291d634c6c6d3d6e4302f75259525721c2aab2edfee6f69405a36ecf2b683

Observation c90c83f6-52ab-4545-84d1-3304d47f63c3 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.819840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.819840Z digest=sha256:bc9947abfa39381ce4233f911b67f1f94f5afae1d88d8119ffcab8215318fe98

Observation 5a38610f-7acc-42ac-a56d-c9ab3c9d4908 · outbound

This paper cites Cradle: Empowering Foundation Agents Towards General Computer Control.

lmgame-Bench: How Good are LLMs at Playing Games? Cradle: Empowering Foundation Agents Towards General Computer Control

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.872288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.872288Z digest=sha256:07b7c531ce21d4ce5ddcb9c850cd9ba1e6f173b426af1de48dafa8172faff467

Observation 091a4179-5910-422c-8688-a32d3cb84c84 · outbound

This paper cites In The Twelfth International Conference on Learning Representations.

lmgame-Bench: How Good are LLMs at Playing Games? In The Twelfth International Conference on Learning Representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.941657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.941657Z digest=sha256:6d4bbb7a557df71d69ffb96fc5836291ab5791bc2b4286acb6621954e79cd9d6

Observation e14d8a20-afe8-440c-b5d0-b4b75a82f599 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.038784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.038784Z digest=sha256:996e71c34ca1b8d7c64a59a4ba861b5797f111e0677126680e4bdea1b8158a67

Observation 1f4eaf26-1294-4981-812f-a1f1136a981e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

lmgame-Bench: How Good are LLMs at Playing Games? Measuring Massive Multitask Language Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.115491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.115491Z digest=sha256:e3396f9102f492e2439ff715c77bbb1dc631a8dc6deff4650a42af0054fb2f10

Observation 78264078-8bb5-473a-a8db-9465570d5a75 · outbound

This paper cites Humanity's Last Exam.

lmgame-Bench: How Good are LLMs at Playing Games? Humanity's Last Exam

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.189173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.189173Z digest=sha256:059070c106e916a77b385f9303405dc66e78db9b56a9603c094b8389686b8247

Observation 5215fa44-fcc9-4428-ac26-30c80fcaa06c · outbound

This paper cites https://scale.com/leaderboard/ humanitys_last_examAccessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://scale.com/leaderboard/ humanitys_last_examAccessed: 2025-05-14

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.256804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.256804Z digest=sha256:2ea74adda3f43d277071637036a9792d77968995845bb839cbe3d4634cbe467e

Observation fa21a2a7-b374-4ee5-8416-5e1e6fa90698 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.333175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.333175Z digest=sha256:d6a09c26f50cfb64593b9f15b2ec5edd8fb7354025a79141c06824143f08028c

Observation 05197398-9457-479b-8f9b-c55b05938893 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

lmgame-Bench: How Good are LLMs at Playing Games? GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.405000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.405000Z digest=sha256:1642781c708c14dcf956cf74b14499b223cac9b4d2f0c63735a2c1b4d2a3dc6d

Observation 18401cb3-ae82-436f-a734-1c1b8cf03368 · outbound

This paper cites https://www.vals.ai/benchmarks/ gpqa-05-09-2025Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ gpqa-05-09-2025Accessed: 2025-05-14

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.464341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.464341Z digest=sha256:ccdf386565a63083c64c2bd63c9b310cb731a1f59ee2cc5a474222a8bd0305d6

Observation c0b9f034-2e97-42c1-bb1a-d5525b541192 · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

lmgame-Bench: How Good are LLMs at Playing Games? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.524310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.524310Z digest=sha256:dd0cccdde7b4b911915582caa4fc259b12ee7fc1b5fe5153f5359e8e5ef48c8f

Observation 8cafd74e-8f58-4fde-a158-02ea21ec09c9 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

lmgame-Bench: How Good are LLMs at Playing Games? Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.566161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.566161Z digest=sha256:f2cc96271d2f41c95b6ae70d83bd07330602a8cf228b666b2c0a4f44425db0a6

Observation cce49f77-c6ae-4fb2-aab9-75b4ec1248d6 · outbound

This paper cites https://www.vals.ai/benchmarks/ math500-05-09-2025Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ math500-05-09-2025Accessed: 2025-05-14

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.676392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.676392Z digest=sha256:937c92bc0dab35980c2be90c247a29b7d169b61beafba0e94797740d552490c1

Observation 90b9f4e3-156d-4664-aee1-57fb91077fdf · outbound

This paper cites arXiv preprint arXiv:2410.03131 (2024).

lmgame-Bench: How Good are LLMs at Playing Games? arXiv preprint arXiv:2410.03131 (2024)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.788838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.788838Z digest=sha256:f7cffdba4753380883f88dff2410c324744fabc0298006564b7e333b6c82f78f

Observation fd08dd1e-6fbf-48a1-98b3-78fd6a87369b · outbound

This paper cites https://www.vals.ai/benchmarks/ aime-2025-05-09Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ aime-2025-05-09Accessed: 2025-05-14

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.975697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.975697Z digest=sha256:a06942de8b62d6610b6cc8cdcd14cd05a1bcf865528d120912095aed3234316c

Observation b2b1ae1a-b149-46a6-8cf8-7810df23e33b · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

lmgame-Bench: How Good are LLMs at Playing Games? LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.152304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.152304Z digest=sha256:9e5655b3f046c93f489703755c1d8d133e62adc7458bc437d72f4553e02a7c96

Observation 57ad3dd4-a046-4304-b642-c79511ba15da · outbound

This paper cites https://livebench.ai/#/?Coding=a& Mathematics=a&Data+Analysis=a&Language=a&IF=aAccessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://livebench.ai/#/?Coding=a& Mathematics=a&Data+Analysis=a&Language=a&IF=aAccessed: 2025-05-14

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.352797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.352797Z digest=sha256:28570ed9f2f84f53f646829427b9a1df356931a622f1cd615f6658c4f0335696

Observation 57249563-1ef7-4de2-a8ef-873d95490edd · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

lmgame-Bench: How Good are LLMs at Playing Games? BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.559765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.559765Z digest=sha256:07d492a7f37451e6b56c68fe6d20a8181ee580edf5190738cc3f323bbeba7def

Observation 417548ee-d2f4-4b93-a41b-3768f88ef4a6 · outbound

This paper cites https://aider.chat/docs/leaderboards/ Ac- cessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://aider.chat/docs/leaderboards/ Ac- cessed: 2025-05-14

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.736996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.736996Z digest=sha256:8c2618f952ffde6eff7709d15464972c05598f65fb84deebc58110d84097b681

Observation f74dbb8d-05cd-47d5-805c-429e6fb2df25 · outbound

This paper cites https://bigcode-bench.github.io/ Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://bigcode-bench.github.io/ Accessed: 2025-05-14

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.875924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.875924Z digest=sha256:0b784f21ef3abb094c1a13ce9e7599d0df455e05f4830c0fbdf0350ddbae018c

Observation 5f981890-9fcb-428b-ba47-a854b244afe3 · outbound

This paper cites https://scale.

lmgame-Bench: How Good are LLMs at Playing Games? https://scale

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.977620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.977620Z digest=sha256:caec3e240aa94f069e1310944aea31fe14709b1127ea2fbda8b094df6656bc4e

Observation 1b3e83b5-e1af-4700-83e5-b546bd993e2a · outbound

This paper cites https://lmarena.ai/leaderboard Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://lmarena.ai/leaderboard Accessed: 2025-05-14

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.981839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.981839Z digest=sha256:4f14fe8c7d06fea557386fb34965079d35382b39c961f36c1af6c0a5e9b5ba93

Observation c8b4e03f-6aa8-40d0-9580-048940736906 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

lmgame-Bench: How Good are LLMs at Playing Games? MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.986636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.986636Z digest=sha256:d27aaf3a4267b07645a4f5afa9bcc8f1cc58c7476443dfa3614ee1bdab0fb737

Observation 46b6f099-93db-472a-bbb9-5991bf80b0b2 · outbound

This paper cites https://www.vals.ai/benchmarks/ mmmu-05-09-2025Accessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ mmmu-05-09-2025Accessed: 2025-05-14

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.031667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.031667Z digest=sha256:ac6f4242cd8383a026d297b0f51262f100e020f8f488fc22b048dc092d1d9f5a

Observation 8158aad3-a061-4172-a5e1-392d3e4fa3a4 · outbound

This paper cites MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs.

lmgame-Bench: How Good are LLMs at Playing Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.110876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.110876Z digest=sha256:f8d1150b226668823abc99ce6c124093ae88ce127d17b6b9fdaf4bd4ba3e0116

Observation 3e400887-44cc-4397-99cf-e6dcb36ebaa9 · outbound

This paper cites https://scale.com/leaderboard/ multichallengeAccessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://scale.com/leaderboard/ multichallengeAccessed: 2025-05-14

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.208322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.208322Z digest=sha256:166ebe418669ac198026c10d2a6a7a7e228c1a107cfa92981ebcb1d62994653f

Observation e543a205-cfd9-4a73-8081-570c10c774b0 · outbound

This paper cites https://scale.com/leaderboard/ enigma_evalAccessed: 2025-05-14.

lmgame-Bench: How Good are LLMs at Playing Games? https://scale.com/leaderboard/ enigma_evalAccessed: 2025-05-14

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.299976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.299976Z digest=sha256:289caab232af0e0b52cf86c0a1d3f089fad0814eacdccf07148d455463d22118

Observation d1b43e71-9dfc-4115-b0db-dde5f55ad797 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.362623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.362623Z digest=sha256:3a5a19d9e7da47129433dc0fb37fe0df62a730389689b7a843698fd5d22f6742

Observation 9b944751-7bca-4f85-a541-36dd67dd2b0b · outbound

This paper cites https://github.com/mpSchrader/gym-sokoban (2018).

lmgame-Bench: How Good are LLMs at Playing Games? https://github.com/mpSchrader/gym-sokoban (2018)

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.442781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.442781Z digest=sha256:a40a02b0307f014f1e63ce8aa7594f7f348d3ce6ed03a224c6e674bb8fff59ce

Observation 942416c6-c81b-4b08-a263-8b1b5a540611 · outbound

This paper cites https://github.com/jaybutera/ tetrisRL(2023) GitHub repository.

lmgame-Bench: How Good are LLMs at Playing Games? https://github.com/jaybutera/ tetrisRL(2023) GitHub repository

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.626893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:01.518032Z digest=sha256:90073597b34786344bf299d93bcde1d2f6b5c8824f26d59477c766c2ad27d47d

Observation 086da939-2779-4dbc-9d8a-e528b44e84a9 · outbound

This paper cites Qwen2.5 Technical Report.

lmgame-Bench: How Good are LLMs at Playing Games? Qwen2.5 Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.523920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.523920Z digest=sha256:a442449cb5487579e1750861c8fe9811ffc9f8aa5ff6cf6d5b37bac8441f12e7

Observation cf891336-6b2c-40d8-abd9-19785f67c849 · outbound

This paper cites Communications of the ACM 38(3) (1995) 58–68.

lmgame-Bench: How Good are LLMs at Playing Games? Communications of the ACM 38(3) (1995) 58–68

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.609837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:01.528538Z digest=sha256:4060507051e3cc2b25831d6f923d67b03823af67e93f6cf59be360ce194a5dab

Observation b0d4f444-27e6-4395-a8e0-5ae73574d782 · outbound

This paper cites nature550(7676) (2017) 354–359.

lmgame-Bench: How Good are LLMs at Playing Games? nature550(7676) (2017) 354–359

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.592570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:01.544702Z digest=sha256:d13f963322a5a7d60c63973f3b45c645ec33b798fa7f88e5dcdb6e801e43c4d6

Observation 04da8a84-7fdb-4484-9b9a-ef622d288117 · outbound

This paper cites In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track.

lmgame-Bench: How Good are LLMs at Playing Games? In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.576974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:01.617277Z digest=sha256:7210031dce56c9611d8248994fbeac163bbba2904329c027f38fa91af07c1a72

Observation f9065042-7769-4785-a445-20b67486def2 · outbound

This paper cites Factorio Learning Environment.

lmgame-Bench: How Good are LLMs at Playing Games? Factorio Learning Environment

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.689046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.689046Z digest=sha256:a5a1c83ff9aeded3d36c7146894e0ea54c3d36314e5b413d2d722a852d323ca4

Observation 9a6fa030-734c-47ee-ad2d-e8c8af2764d1 · outbound

This paper cites In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track.

lmgame-Bench: How Good are LLMs at Playing Games? In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.560892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:01.766369Z digest=sha256:f46b3de369f30edb0605d414ae602386daf867906c4e3dc1b5ce1f3c4477ddd0

Observation 16b11839-c133-4a36-bcd0-05ae76aa1262 · outbound

This paper cites TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning.

lmgame-Bench: How Good are LLMs at Playing Games? TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.826741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.826741Z digest=sha256:86c8f145e35860d81e6a6279d086d9c748b96bbb9ba6ec400e369f666090fa5b

Observation ef4dd5ca-d0f2-4830-a96b-26312327dea1 · outbound

This paper cites GameEval: Evaluating LLMs on Conversational Games.

lmgame-Bench: How Good are LLMs at Playing Games? GameEval: Evaluating LLMs on Conversational Games

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.892670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.892670Z digest=sha256:788225688ad2a92c1109839e51bb68e129f1d212624bc891c057120b800cc966

Observation ff01265b-d150-4cf4-bf94-37cd01ea52a9 · outbound

This paper cites GameArena: Evaluating LLM Reasoning through Live Computer Games.

lmgame-Bench: How Good are LLMs at Playing Games? GameArena: Evaluating LLM Reasoning through Live Computer Games

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.953231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.953231Z digest=sha256:5b8133e35ba9f52f20e6594669e939c1b3d10a98b1eca3ea6362a0eef09b560f

Observation 6a8c9fc0-bf4f-4d18-a4ec-336252fdc380 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

lmgame-Bench: How Good are LLMs at Playing Games? SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.019566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.019566Z digest=sha256:e163db42a62808a997ab950d6a85e0fb8a9c8af2ed5f8b988c4fc9e925e5efca

Observation 93c0d409-ab27-49cb-8248-b6a7e5e39c44 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

lmgame-Bench: How Good are LLMs at Playing Games? WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.059918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.059918Z digest=sha256:53668c64f59dbdb6fe97d63167b50394544f6aabe4637642897f92c289a3554b

Observation 71003540-eb94-4f0a-8628-80fdb703865a · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

lmgame-Bench: How Good are LLMs at Playing Games? WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.065525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.065525Z digest=sha256:73c8ff5f170c3112224e39da9dc6353b27c73ad0ef6b1c5631435770da7e15b8

Observation fbbce248-f0d4-44ce-9218-ff4a600c30a9 · outbound

This paper cites AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents.

lmgame-Bench: How Good are LLMs at Playing Games? AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.082501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.082501Z digest=sha256:dd16d0f5c21d9d7f47dd7932f046c226701688a1fbd2dc8b7864b1b2d1656fc1

Observation 55c10a9a-0b26-4173-92d4-352510af82cc · outbound

This paper cites Advances in Neural Information Processing Systems37(2024) 52040–52094.

lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems37(2024) 52040–52094

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.545710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.131440Z digest=sha256:e35e655bcf39dd21f9c824e7fb1a79732e437490d73f53d4522282f2bfbcf1e3

Observation 23603768-7adc-4214-82f5-134ebcf65371 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:27:04.528964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.208401Z digest=sha256:a7312f14b66ef062cdf4995a520e0a25bcfbcd963dfb4193a04db561aca65788

Observation ab6ef914-7e02-48fc-ac4c-3a0923f26339 · outbound

This paper cites In The Twelfth International Conference on Learning Representations.

lmgame-Bench: How Good are LLMs at Playing Games? In The Twelfth International Conference on Learning Representations

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.513351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.293383Z digest=sha256:fe4a6d0a3fe9aa8a11bca093d0c0b8cbec4a66a7f6cb9f7b2e97b940151c87f1

Observation ac079a74-203a-44ca-a45c-7e4e7b81e5a3 · outbound

This paper cites Advances in neural information processing systems37(2024) 110935–110971.

lmgame-Bench: How Good are LLMs at Playing Games? Advances in neural information processing systems37(2024) 110935–110971

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.497010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.328494Z digest=sha256:2958453019ae14250d4ac809e123357c28d161736bedb30ca76024ebf44b87a6

Observation dd33cdcf-25ab-4ac1-8e0a-508f1a56a9f0 · outbound

This paper cites Proximal Policy Optimization Algorithms.

lmgame-Bench: How Good are LLMs at Playing Games? Proximal Policy Optimization Algorithms

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.348919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.348919Z digest=sha256:9951abf74ee7bcc4db3559acffa9228895a9aea528028c9d69ff1c854b13d754

Observation f23d2550-8302-4b52-9bb2-70b39c762a73 · outbound

This paper cites Science10(3) (1995) 237–304.

lmgame-Bench: How Good are LLMs at Playing Games? Science10(3) (1995) 237–304

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.478783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.397270Z digest=sha256:4e2fa731475d9a4ed8a9574e78914fac902cc4c3805f03e039bbbee4b849e4fe

Observation d17c69fb-0447-4821-94ee-8b7e972aff5c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

lmgame-Bench: How Good are LLMs at Playing Games? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.445271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.445271Z digest=sha256:fde628613a91687346cf6cc15c2fa974932d9e2fe351b275caaee10828cb8a23

Observation bc3d9d72-699d-4e55-a952-6d4994acc0ff · outbound

This paper cites Advances in Neural Information Processing Systems36(2023) 38975–38987.

lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems36(2023) 38975–38987

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.461822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.509094Z digest=sha256:b9926f812fea9541968b97a632efd9b18c8bc5fe72c4008cd5f99f3c3796f43d

Observation 87ddd352-621e-4aab-a607-a901a790b5ca · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

lmgame-Bench: How Good are LLMs at Playing Games? Training Verifiers to Solve Math Word Problems

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.575808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.575808Z digest=sha256:acd6a001e47e2c2e2c8014b9537c288e487a926bcf809cdef164ee4de761cfec

Observation 102da256-cfb9-408f-8b4a-80014645a6b9 · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems36(2024)

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.444914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.602474Z digest=sha256:215eb0061d273f892ac7d09dd81a4091a9834b0284e390dc123567f416a6c2f9

Observation 25bdbe78-156d-4464-be03-03384ac850ac · outbound

This paper cites Advances in Neural Information Processing Systems 35(2022) 20744–20757.

lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems 35(2022) 20744–20757

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.426651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.606960Z digest=sha256:7113ba5c80398b6d83f013c45316f147346cd6371b42a33c21ed7365228f43bd

Observation 82404a63-dd5e-4dc3-bfea-47ec0b60f14c · outbound

This paper cites Biometrika30(1/2) (1938) 81–93.

lmgame-Bench: How Good are LLMs at Playing Games? Biometrika30(1/2) (1938) 81–93

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.409726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.679885Z digest=sha256:9d9f95ec06d3cbfe40ab374079a6ca4e5af9a375fa506d7bcd0c6a434782ab96

Observation 49d0ee1f-ec40-4b19-b216-87c56e6c6c61 · outbound

This paper cites In Proceedings of the 19th international conference on World wide web, ACM (2010) 577–586.

lmgame-Bench: How Good are LLMs at Playing Games? In Proceedings of the 19th international conference on World wide web, ACM (2010) 577–586

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.392662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.740913Z digest=sha256:363dac534be47143e220663b203c0097a95f851b4bc56864e2845d70ce9c31e2

Observation 4f03dc3a-ea7e-47b1-8691-d7281719da5a · outbound

This paper cites Educational Researcher5(10) (1976) 3–8.

lmgame-Bench: How Good are LLMs at Playing Games? Educational Researcher5(10) (1976) 3–8

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.376853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.809560Z digest=sha256:b5eb295f701a0fc75fc3fc19ed8a19177f68cda4d8b7f4860cc0001c33b65eb3

Observation fb942280-1434-4e92-8e42-927df2631386 · outbound

This paper cites Block 1 is on top of block 3, block 3 is on top of block 2, and block 2 is on the table.

lmgame-Bench: How Good are LLMs at Playing Games? Block 1 is on top of block 3, block 3 is on top of block 2, and block 2 is on the table

Reference 91

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:27:04.361377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:02.852880Z digest=sha256:c03c8f1521f0b486bba91de2f04d873baa82eb858faa429a10e6211467d3cdc8

Observation 616a293f-add6-4048-8061-6e2fee8b6378 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.911988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.911988Z digest=sha256:7a3702f855a1260a4ea4d1c9a136140fb6b9fd2ee89dd18545bd0cf3f31bb008

Observation 9dda3d17-5a70-466d-8d56-55981e5b0ef1 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.950794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.950794Z digest=sha256:1788f8af0a39224ffc5189e8184f8547e3dcd40c5a1ad2388e2499e39f94dc85

Observation e1830c7f-0d00-443b-9b70-32409578e2b0 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.032310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.032310Z digest=sha256:c08c389f796a01642f837f4e9199c6ef16835d61e8a24c19db74194bb2d6462e

Observation 56d4f2d6-0d9a-4a9e-8129-55c3b56a358f · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.119984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.119984Z digest=sha256:62a9b73212efa143708d0f7b3c46dd420623a5ac5438d4639f18662dc3e81c02

Observation fff924be-933e-4c74-bb34-f0059920b182 · outbound

This paper cites up", "down.

lmgame-Bench: How Good are LLMs at Playing Games? up", "down

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.304373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:27:03.163759Z digest=sha256:384c3ea4847094b525af9570905cd9fe2fa32307f293e1c8d8b377004367931e

Observation ce3e4880-c2cc-45a0-ae20-063acc143f72 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.171436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.171436Z digest=sha256:16c67621cfc69a9ab38945bac124aac8ba0e0b95511c058473479dc2d066a263

Observation f08e1eca-d190-4a28-b8c3-50d5a956d03c · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.177177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.177177Z digest=sha256:5c3049ad243181645f374cf48a2403fff400967bac7663c88dce19622a7e6cff

Observation 1025f4c9-f81b-4b4d-b541-07f78456c050 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.182638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.182638Z digest=sha256:4117c73143ae4c4690dcd8e77721688c55bf7bdc57d3d454dd7ad6891d82ca35

Observation c58fb195-9556-4c8d-ae4b-178daf346b75 · outbound

This paper cites an unresolved cited work.

lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:03.187728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:03.187728Z digest=sha256:04d5141863bc4878a83985980f529e39577090c91776f71ee4f3c4f55cfe5f92

Pith citing papers

Observation 6eff8939-1f26-465a-9aab-d505c04f2523 · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers lmgame-Bench: How Good are LLMs at Playing Games?

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:33.015304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:33.015304Z digest=sha256:5e53663f424cf07145f77f73f2b8857927b151a36172ffcd74a8fc7b99033848

Observation 6b6def62-bd00-4966-905d-2beeba9e4b8f · inbound

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play cites this paper.

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play lmgame-Bench: How Good are LLMs at Playing Games?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:33:42.500139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:33:42.500139Z digest=sha256:ebe60848a1c9bc35d0d3bbb6b2da85839868652632396558556f727951aa3159

Observation c87a4664-3ed6-421e-bf53-762d56b3599c · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning lmgame-Bench: How Good are LLMs at Playing Games?

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.184183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:6ad254b7cabdeba95d871c8da460d49f8614e162fa062640b16992ecb8c28b7c

Observation 27ec31b8-d11b-4952-889c-a66ceb5987ad · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models lmgame-Bench: How Good are LLMs at Playing Games?

Reference 203

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:05:31.526497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:e7c888ba7e05858ecc0a3badca5729765234f48ce168af8786b254e399c59c84

Observation f8567ab9-4edc-49b5-a74d-036f1e741261 · inbound

Gym-V: A Unified Vision Environment System for Agentic Vision Research cites this paper.

Gym-V: A Unified Vision Environment System for Agentic Vision Research lmgame-Bench: How Good are LLMs at Playing Games?

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:05:26.044132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T10:05:09.049846Z digest=sha256:cab3b5f6c5a8531b8311570eda7a2593e8f4e465699304055aa6322107229c36

Observation 5f80c1c4-4c08-452c-a764-51911f301c29 · inbound

TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs cites this paper.

TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs lmgame-Bench: How Good are LLMs at Playing Games?

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:16:06.152542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:04:57.744754Z digest=sha256:da9d4f7bf448cdc5ba71c9313f24dd87bdb42757d4e9788987839e707ad4b094

Observation e18a66e5-bba2-4612-b96a-11391a6046d7 · inbound

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks cites this paper.

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks lmgame-Bench: How Good are LLMs at Playing Games?

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:14:46.428384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T00:14:07.017420Z digest=sha256:3d9e9e0c0db6ffc80bf6720b5d0e3bdee6bb6b7fefd0fa054772331183c83484

Observation 1fb4d952-c576-432d-96bc-4165ca85794e · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning lmgame-Bench: How Good are LLMs at Playing Games?

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:09.364638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:c35a21404a3f282c3662a45656c57ecf000f8444ab0174d6af718a749214c628

Observation 70f766f4-694a-44b7-b27c-a41a16861d3c · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:07:51.419743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T19:05:36.511150Z digest=sha256:e8672c3ef875bc898a4fd08f408257d8e0cdaabf211c5328bd6c8bd3ec2bca65

Observation 5d47647a-68a4-4f89-984d-0df532ab97c7 · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:59:48.141281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T05:59:44.669877Z digest=sha256:3c5314a92be0b61b6e73eee29e6aa1b00f3eb578eb5d45a90291f049a4ff5158

Observation 1d368f40-3266-41ea-8cdd-8b9fb6fa95df · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.559898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:7d76acee8696e5faded12fe4a6a72e1924e9bd7d6d93a7d81f6fac94e7ecd33f

Observation cfd7b7b2-11a5-41eb-8577-10df4fcd792b · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games lmgame-Bench: How Good are LLMs at Playing Games?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:17.261880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T12:24:06.062957Z digest=sha256:2d2d3af7b4a5577a47183479495da0875c16014d1bad58d1bf17efb0eadcfbf5

Observation 2049d76d-8774-4384-a78f-ba5c6d631ed4 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games lmgame-Bench: How Good are LLMs at Playing Games?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:45:23.221897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T05:45:04.573722Z digest=sha256:5253154836eda1f616e594d1fc782d617c2b4808d83c60b87c2c6926fef769a4

Observation cb5202f0-d7ec-4fae-b9ff-6d58f9651a63 · inbound

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs cites this paper.

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs lmgame-Bench: How Good are LLMs at Playing Games?

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:13.196706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T07:15:27.939886Z digest=sha256:ecf55f6494da18290c0af6ebb63a4485072abdcbf42b214635a0fc32a011d26c

Observation 23f86099-b0a4-427e-800a-d1b8d9bdeb5f · inbound

PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game? cites this paper.

PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game? lmgame-Bench: How Good are LLMs at Playing Games?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.593221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T07:21:49.763994Z digest=sha256:1f1777a046fb8f5b63ef85e2f4e9ea0e4d2b3e0e1e7a5abde1f5d7d43bf5e8c0

Observation cad8132a-3e76-4034-aed2-e4abc28a1d66 · inbound

Robots Need More than VLA and World Models cites this paper.

Robots Need More than VLA and World Models lmgame-Bench: How Good are LLMs at Playing Games?

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:11:28.899789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T01:01:33.530167Z digest=sha256:44a9250a8c9137bdf3922a122d8c9b8e0a06be19544e3ef86f3879b08f28534a

Observation f4c3ebe9-e250-4a38-9fe2-063b44fa6a8c · inbound

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics cites this paper.

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics lmgame-Bench: How Good are LLMs at Playing Games?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:29.947294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T16:50:36.194650Z digest=sha256:9d7379378ee89c349f016b6d9b8b137dada5e16a61c3a1d75fbd3761709319b5

Observation 09f90fa0-57f8-4c7f-9529-c4a8dd5f4cc1 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application lmgame-Bench: How Good are LLMs at Playing Games?

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:50:48.495612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:5045484f427bed043d80818811fe4bfc88ab8a359da35235f563738f43226eed

Observation 0be5c38e-8540-42a0-be0c-76e3c213b0e2 · inbound

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models cites this paper.

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models lmgame-Bench: How Good are LLMs at Playing Games?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.299890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T21:02:26.081166Z digest=sha256:3fc88accb7ab36960850fe5641b4a9901a424a91542eba2d6e68a0e3729b2a15

Observation 914f01fc-110b-4bcd-9f5a-dbb5aad4ede8 · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games lmgame-Bench: How Good are LLMs at Playing Games?

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:19:13.718478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:03f7463e46f136e13cbff04c4b8976b69f106c0d1eaaba7820d8ded6da73e391

Observation 0c2486de-cfd0-4b84-97a8-f3b4a71f6f00 · inbound

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents cites this paper.

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:46:11.447161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T21:59:25.449570Z digest=sha256:a29c5ab2c0a9ec3f7a37256e9a0f775ca5821e14c10af4c87cfa3a74019765d3

Observation c51221f7-15e7-4980-ac29-d6e0cb128700 · inbound

CAST: Game Solvers as Turn-Level Teachers for LLM Agents cites this paper.

CAST: Game Solvers as Turn-Level Teachers for LLM Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T02:58:29.283205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:58:29.283205Z digest=sha256:0bc119051d6c946f366124935fbb742a40b88f756d0fc670114200f14effa8d2