Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:27:03.187728Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 115 outbound references and 22 inbound Pith citation observations for arXiv:2505.15146.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:27:03.187728Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.015304Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
100 of 115 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation a7e70288-ebc0-4e09-82a6-f0ea6f228c64 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? OpenAI Gym
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131c4169-85b2-400a-865b-adda2c05762f · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba224b0f-9985-4a21-bb0e-6fa7100a6c5c · outbound
lmgame-Bench: How Good are LLMs at Playing Games? RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a9e272-3755-4f82-8ebb-4c5d0b603de9 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d9d2fe6-ad2e-4a8c-ae80-807da734ae0b · outbound
lmgame-Bench: How Good are LLMs at Playing Games? In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., eds.: Advances in Neural Information Processing Systems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d0d7fbc-84b1-43ad-a74d-6b37e571ae55 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation babc8d0b-adff-4d21-9ae3-ee3d6a759645 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57530e12-cda5-41fe-bf7f-6574601f213e · outbound
lmgame-Bench: How Good are LLMs at Playing Games? EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ca5e66-1c30-491b-9df9-19c7522ef01d · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc84ddec-7530-45b2-aa78-08091fdbc4a6 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e38bd61-6f22-4a25-8709-7c5227b6d909 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb50f57-f0b3-4573-880d-350a08c99f3b · outbound
lmgame-Bench: How Good are LLMs at Playing Games? SmartPlay: A Benchmark for LLMs as Intelligent Agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa6771e-9efd-40a8-8ccd-f8cd596c8f53 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? IEEE Transactions on Games11(3) (2019) 195–202
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97837f3f-1fdd-4b8d-9c47-82e019956ede · outbound
lmgame-Bench: How Good are LLMs at Playing Games? AI Magazine22(2) (2001) 15–25
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14041c05-4acb-4b6b-bb51-360566b36a2b · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2607c706-b810-4f3e-a02e-76c5f1a9e56e · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42618b84-232d-4f90-b0d7-8e954d3735c2 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e90353b-7af8-4f73-8b05-213dc2a49d13 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c06fd2b-4bce-456f-ae46-6b275c790620 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? OpenAI o1 System Card
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2613515-e9ef-4538-b0cb-848ce7429f6a · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f282f996-6b9a-4e1f-8578-ad4a26aa3d85 · outbound
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed9dfa1-d833-40fa-af0e-82ebf70154f0 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Applied cognitive psychology31(4) (2017) 438–445
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3791a52-c555-4ec5-bb2f-a3ea2101db68 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? In International Computing and Combinatorics Conference (COCOON)
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a661cdd9-5b21-433f-b165-7d7818203605 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458995a7-b87c-48cf-892d-495c8af4cd2b · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d50081-a23e-4b11-ad68-4a9665b646da · outbound
lmgame-Bench: How Good are LLMs at Playing Games? arXiv (2024)
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eadfffa4-9db8-4542-9f36-cc139a0416a3 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? arXiv (2025)
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9973ff4-17fc-4c65-a76e-ebac676753b3 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d32840-27ff-4a3a-95c4-4e9206bb84f4 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Computational Geometry 13(4) (1999) 215–228
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 315386b0-a926-4c90-b7d8-da2b71f4b8b2 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? A Simple Family of Analytical Trumpet Slices of the Schwarzschild Spacetime
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df103d95-f26f-4379-a55b-5e402af91911 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6f11591-6ffc-4ad2-a2a3-a1758280f791 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 779eca80-188b-4c50-852e-178d48195594 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90c83f6-52ab-4545-84d1-3304d47f63c3 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a38610f-7acc-42ac-a56d-c9ab3c9d4908 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Cradle: Empowering Foundation Agents Towards General Computer Control
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091a4179-5910-422c-8688-a32d3cb84c84 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? In The Twelfth International Conference on Learning Representations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14d8a20-afe8-440c-b5d0-b4b75a82f599 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f4eaf26-1294-4981-812f-a1f1136a981e · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Measuring Massive Multitask Language Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78264078-8bb5-473a-a8db-9465570d5a75 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Humanity's Last Exam
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5215fa44-fcc9-4428-ac26-30c80fcaa06c · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://scale.com/leaderboard/ humanitys_last_examAccessed: 2025-05-14
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa21a2a7-b374-4ee5-8416-5e1e6fa90698 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05197398-9457-479b-8f9b-c55b05938893 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18401cb3-ae82-436f-a734-1c1b8cf03368 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ gpqa-05-09-2025Accessed: 2025-05-14
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b9f034-2e97-42c1-bb1a-d5525b541192 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cafd74e-8f58-4fde-a158-02ea21ec09c9 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cce49f77-c6ae-4fb2-aab9-75b4ec1248d6 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ math500-05-09-2025Accessed: 2025-05-14
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b9f4e3-156d-4664-aee1-57fb91077fdf · outbound
lmgame-Bench: How Good are LLMs at Playing Games? arXiv preprint arXiv:2410.03131 (2024)
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd08dd1e-6fbf-48a1-98b3-78fd6a87369b · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ aime-2025-05-09Accessed: 2025-05-14
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2b1ae1a-b149-46a6-8cf8-7810df23e33b · outbound
lmgame-Bench: How Good are LLMs at Playing Games? LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ad3dd4-a046-4304-b642-c79511ba15da · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://livebench.ai/#/?Coding=a& Mathematics=a&Data+Analysis=a&Language=a&IF=aAccessed: 2025-05-14
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57249563-1ef7-4de2-a8ef-873d95490edd · outbound
lmgame-Bench: How Good are LLMs at Playing Games? BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417548ee-d2f4-4b93-a41b-3768f88ef4a6 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://aider.chat/docs/leaderboards/ Ac- cessed: 2025-05-14
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f74dbb8d-05cd-47d5-805c-429e6fb2df25 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://bigcode-bench.github.io/ Accessed: 2025-05-14
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f981890-9fcb-428b-ba47-a854b244afe3 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://scale
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b3e83b5-e1af-4700-83e5-b546bd993e2a · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://lmarena.ai/leaderboard Accessed: 2025-05-14
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b4e03f-6aa8-40d0-9580-048940736906 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b6f099-93db-472a-bbb9-5991bf80b0b2 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://www.vals.ai/benchmarks/ mmmu-05-09-2025Accessed: 2025-05-14
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8158aad3-a061-4172-a5e1-392d3e4fa3a4 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e400887-44cc-4397-99cf-e6dcb36ebaa9 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://scale.com/leaderboard/ multichallengeAccessed: 2025-05-14
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e543a205-cfd9-4a73-8081-570c10c774b0 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://scale.com/leaderboard/ enigma_evalAccessed: 2025-05-14
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b43e71-9dfc-4115-b0db-dde5f55ad797 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b944751-7bca-4f85-a541-36dd67dd2b0b · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://github.com/mpSchrader/gym-sokoban (2018)
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 942416c6-c81b-4b08-a263-8b1b5a540611 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? https://github.com/jaybutera/ tetrisRL(2023) GitHub repository
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 086da939-2779-4dbc-9d8a-e528b44e84a9 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Qwen2.5 Technical Report
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf891336-6b2c-40d8-abd9-19785f67c849 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Communications of the ACM 38(3) (1995) 58–68
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b0d4f444-27e6-4395-a8e0-5ae73574d782 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? nature550(7676) (2017) 354–359
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 04da8a84-7fdb-4484-9b9a-ef622d288117 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f9065042-7769-4785-a445-20b67486def2 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Factorio Learning Environment
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a6fa030-734c-47ee-ad2d-e8c8af2764d1 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16b11839-c133-4a36-bcd0-05ae76aa1262 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef4dd5ca-d0f2-4830-a96b-26312327dea1 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? GameEval: Evaluating LLMs on Conversational Games
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff01265b-d150-4cf4-bf94-37cd01ea52a9 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? GameArena: Evaluating LLM Reasoning through Live Computer Games
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a8c9fc0-bf4f-4d18-a4ec-336252fdc380 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c0d409-ab27-49cb-8248-b6a7e5e39c44 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71003540-eb94-4f0a-8628-80fdb703865a · outbound
lmgame-Bench: How Good are LLMs at Playing Games? WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbbce248-f0d4-44ce-9218-ff4a600c30a9 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c10a9a-0b26-4173-92d4-352510af82cc · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems37(2024) 52040–52094
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 23603768-7adc-4214-82f5-134ebcf65371 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ab6ef914-7e02-48fc-ac4c-3a0923f26339 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? In The Twelfth International Conference on Learning Representations
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ac079a74-203a-44ca-a45c-7e4e7b81e5a3 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Advances in neural information processing systems37(2024) 110935–110971
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dd33cdcf-25ab-4ac1-8e0a-508f1a56a9f0 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Proximal Policy Optimization Algorithms
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f23d2550-8302-4b52-9bb2-70b39c762a73 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Science10(3) (1995) 237–304
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d17c69fb-0447-4821-94ee-8b7e972aff5c · outbound
lmgame-Bench: How Good are LLMs at Playing Games? DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc3d9d72-699d-4e55-a952-6d4994acc0ff · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems36(2023) 38975–38987
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 87ddd352-621e-4aab-a607-a901a790b5ca · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Training Verifiers to Solve Math Word Problems
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 102da256-cfb9-408f-8b4a-80014645a6b9 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems36(2024)
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 25bdbe78-156d-4464-be03-03384ac850ac · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Advances in Neural Information Processing Systems 35(2022) 20744–20757
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82404a63-dd5e-4dc3-bfea-47ec0b60f14c · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Biometrika30(1/2) (1938) 81–93
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49d0ee1f-ec40-4b19-b216-87c56e6c6c61 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? In Proceedings of the 19th international conference on World wide web, ACM (2010) 577–586
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4f03dc3a-ea7e-47b1-8691-d7281719da5a · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Educational Researcher5(10) (1976) 3–8
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb942280-1434-4e92-8e42-927df2631386 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Block 1 is on top of block 3, block 3 is on top of block 2, and block 2 is on the table
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 616a293f-add6-4048-8061-6e2fee8b6378 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dda3d17-5a70-466d-8d56-55981e5b0ef1 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1830c7f-0d00-443b-9b70-32409578e2b0 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d4f2d6-0d9a-4a9e-8129-55c3b56a358f · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff924be-933e-4c74-bb34-f0059920b182 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? up", "down
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ce3e4880-c2cc-45a0-ae20-063acc143f72 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f08e1eca-d190-4a28-b8c3-50d5a956d03c · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1025f4c9-f81b-4b4d-b541-07f78456c050 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c58fb195-9556-4c8d-ae4b-178daf346b75 · outbound
lmgame-Bench: How Good are LLMs at Playing Games? Unresolved cited work
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eff8939-1f26-465a-9aab-d505c04f2523 · inbound
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers lmgame-Bench: How Good are LLMs at Playing Games?
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b6def62-bd00-4966-905d-2beeba9e4b8f · inbound
Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play lmgame-Bench: How Good are LLMs at Playing Games?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c87a4664-3ed6-421e-bf53-762d56b3599c · inbound
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning lmgame-Bench: How Good are LLMs at Playing Games?
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 27ec31b8-d11b-4952-889c-a66ceb5987ad · inbound
A Survey of Reinforcement Learning for Large Reasoning Models lmgame-Bench: How Good are LLMs at Playing Games?
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f8567ab9-4edc-49b5-a74d-036f1e741261 · inbound
Gym-V: A Unified Vision Environment System for Agentic Vision Research lmgame-Bench: How Good are LLMs at Playing Games?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f80c1c4-4c08-452c-a764-51911f301c29 · inbound
TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs lmgame-Bench: How Good are LLMs at Playing Games?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e18a66e5-bba2-4612-b96a-11391a6046d7 · inbound
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks lmgame-Bench: How Good are LLMs at Playing Games?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1fb4d952-c576-432d-96bc-4165ca85794e · inbound
Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning lmgame-Bench: How Good are LLMs at Playing Games?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 70f766f4-694a-44b7-b27c-a41a16861d3c · inbound
MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d47647a-68a4-4f89-984d-0df532ab97c7 · inbound
MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1d368f40-3266-41ea-8cdd-8b9fb6fa95df · inbound
MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cfd7b7b2-11a5-41eb-8577-10df4fcd792b · inbound
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games lmgame-Bench: How Good are LLMs at Playing Games?
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2049d76d-8774-4384-a78f-ba5c6d631ed4 · inbound
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games lmgame-Bench: How Good are LLMs at Playing Games?
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb5202f0-d7ec-4fae-b9ff-6d58f9651a63 · inbound
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs lmgame-Bench: How Good are LLMs at Playing Games?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 23f86099-b0a4-427e-800a-d1b8d9bdeb5f · inbound
PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game? lmgame-Bench: How Good are LLMs at Playing Games?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cad8132a-3e76-4034-aed2-e4abc28a1d66 · inbound
Robots Need More than VLA and World Models lmgame-Bench: How Good are LLMs at Playing Games?
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f4c3ebe9-e250-4a38-9fe2-063b44fa6a8c · inbound
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics lmgame-Bench: How Good are LLMs at Playing Games?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 09f90fa0-57f8-4c7f-9529-c4a8dd5f4cc1 · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application lmgame-Bench: How Good are LLMs at Playing Games?
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0be5c38e-8540-42a0-be0c-76e3c213b0e2 · inbound
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models lmgame-Bench: How Good are LLMs at Playing Games?
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 914f01fc-110b-4bcd-9f5a-dbb5aad4ede8 · inbound
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games lmgame-Bench: How Good are LLMs at Playing Games?
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c2486de-cfd0-4b84-97a8-f3b4a71f6f00 · inbound
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents lmgame-Bench: How Good are LLMs at Playing Games?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c51221f7-15e7-4980-ac29-d6e0cb128700 · inbound
CAST: Game Solvers as Turn-Level Teachers for LLM Agents lmgame-Bench: How Good are LLMs at Playing Games?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.