Pith. sign in

Paper Citation Record · LEDGER

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

As of 18 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2508.21365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21365 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:27.591640Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:01.106317Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact6
  • verified fuzzy6
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 3462f19f-6220-49a1-a503-049fa24058ce · outbound

This paper cites Cause and Effect: Can Large Language Models Truly Understand Causality?.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Cause and Effect: Can Large Language Models Truly Understand Causality?

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:23:28.247417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:23.843562Z digest=sha256:4dd374296d7d3390873eb1f2561d5d0841ed2ec8f856d679bef092ef8980a66d

Observation 6039e9bf-ca8c-441f-97b9-581a2cd2d618 · outbound

This paper cites an unresolved cited work.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:23.929458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:23.929458Z digest=sha256:77840cf740baf23c6e7349b6afba7d859b1c144ed9507b9f6917406304214c06

Observation 9b2c6e68-646a-43be-92f7-971cf3f8ff64 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.020818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.020818Z digest=sha256:40be0a7926d7e7cf602bb241594d8ad1c9506ec381e6cafdb6c316b09a9ed00c

Observation 65eeae8c-36e2-48ce-8939-8ceb51a810c7 · outbound

This paper cites MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.072935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.072935Z digest=sha256:08651a32745e38f28f8e42c5b0ca35791737b1bb21a4f94f3faa39f6105787ed

Observation 7a428987-6583-4222-9603-fb9898fd35bc · outbound

This paper cites Bayeschess: A computer chess program based on bayesian networks.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Bayeschess: A computer chess program based on bayesian networks

Reference 5

Resolution
verified exact
doi, observed 2026-08-05T14:23:28.080495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.156489Z digest=sha256:2f12d7e2a41bdbd0c4905eacef43bdb1af1722ad8fc2415b689ab17fd5006ef0

Observation d4b144ae-8637-4860-883c-01ca553fd71a · outbound

This paper cites Font and Tobias Mahlmann.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Font and Tobias Mahlmann

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T14:23:29.926244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.207525Z digest=sha256:b58b7d4cfea499398dc0217a8d35388459994cf9aae82843d59689f94e035d85

Observation ce16c2fa-817c-4f36-9a6e-3bc437a6a1b5 · outbound

This paper cites Enabling self-improving agents to learn at test time with human-in-the-loop guidance.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Enabling self-improving agents to learn at test time with human-in-the-loop guidance

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-05T14:23:29.665817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.271531Z digest=sha256:54e90ed1f04e0da212a535ec70a3125985aa6c21bd7967d2efb4481a321aaebb

Observation 7bb921e6-e04c-48b4-83d7-f04eef122fb0 · outbound

This paper cites Measuring massive multitask language understanding.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Measuring massive multitask language understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.347757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.347757Z digest=sha256:ea4813c72df73abe2dd2da3c8b4cf9629b71c012af53a49bd30e28c202c4fa7f

Observation 712a2b16-a84f-48c1-95db-a3c78fcec02a · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.434666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.434666Z digest=sha256:ecfed9a7fae58bdba93b5dad7ff2e840f9139148935080ad5feafc5a5cafc42a

Observation 4d97786b-a523-4a9f-a8cb-623b7b38872f · outbound

This paper cites A Survey on Large Language Model-Based Game Agents.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models A Survey on Large Language Model-Based Game Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.473717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.473717Z digest=sha256:ff92959c372b0e8539c0b66df60cbd4be387668d23f7826b8719b046b3afb0f9

Observation 37e15f8e-9e20-42f3-b515-e38236208f3f · outbound

This paper cites PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.541597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.541597Z digest=sha256:57d6877f27c7990fc312d06b0f883a5a0127ca2dfa83d69799b39564383b8ec5

Observation b120537f-e788-4299-b64b-8bc3f74f543e · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.920211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.602432Z digest=sha256:1008acb329d1cc931d15d0314d7f7773edac98ed0ac5053d0862c07b81f62bc4

Observation 0a47fa39-e61c-4bd8-90a1-6ee078b5988d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.667850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.667850Z digest=sha256:bb45ea5d4df2eadc58d67ff08c244989039735c72ed8a1c9a5ab44685a30538b

Observation b8d24579-c921-478d-a8ca-613a937f29ae · outbound

This paper cites Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:23:29.386562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.726570Z digest=sha256:1dcc02f428f95f0cdf611660eebe94b657c670f6b852d66608564410d6c7b90f

Observation a7266d34-80bb-40ec-a956-b4b069ea27d6 · outbound

This paper cites School chinese benchmark, 2018.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models School chinese benchmark, 2018

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.727816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.798984Z digest=sha256:8f02df4bb63f9675b5decc4d8a9550dd413ba380f0cefe52edfb83206bf29475

Observation 4d0d153b-c670-4065-85c6-183d48962969 · outbound

This paper cites CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.911042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.911042Z digest=sha256:d9a5af0f4e8c87ffdcb7b212919f0dcc40bd3ee97f5337daaf9018d6f1cb49b6

Observation d0e6f647-06ac-422b-be93-43d35f0c9b2e · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Simpo: Simple preference optimization with a reference-free reward

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.975898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.975898Z digest=sha256:2bdb24aef943deca97b4e6ed5adef95f2160e79f2769efc035213af71c7d0ae1

Observation 8fe91552-3333-4fac-b23d-823e54708096 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Playing Atari with Deep Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.071210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.071210Z digest=sha256:9a67a80dd4ac79c3bee45dbb69c5b67946858f68794953bc7bf100a129c87e39

Observation 2ad2ca66-1586-4690-931c-d94db8641ab5 · outbound

This paper cites Creating pro-level AI for a real-time fighting game using deep reinforcement learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Creating pro-level AI for a real-time fighting game using deep reinforcement learning

Reference 19

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T14:23:29.160079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:25.134829Z digest=sha256:77eabdc1275fd73673868477165cf1e92cf848203db0460c33d02cbdcf8e1679

Observation 30ed4bd9-ea79-4473-95b0-6d637aab7769 · outbound

This paper cites Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, et al.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.589098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:25.311280Z digest=sha256:11984ea0253c33c6bf0b51cef65f46d7bcf0492ad567bf26d00752326956837f

Observation cf1e4c3f-0971-4329-965f-2e2f5a9a7df6 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Manning, Stefano Ermon, and Chelsea Finn

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.379171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.379171Z digest=sha256:465cc4a1df33ddf15c6e2058c94434b30e7e85621c0db41aa076a5dfa926e4b3

Observation 95e7b3ed-3fcd-47db-89ce-a6f1e284b427 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.463390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.463390Z digest=sha256:afce1db667e6f8293050b26957cf1c8142404a433cd7f30bc23601708701a02d

Observation 6ad11b99-f8f3-4a63-b322-a6221a23776a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.587387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.587387Z digest=sha256:2aab8acb40a3d7cb8360b7507fcd39b8f2cf3ac82fe1dcb4d1cb92cd533399c7

Observation 5678c3e7-8d6f-494f-b41c-ad52a791efca · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.667622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.667622Z digest=sha256:19374d60711c9e65f55943f59636db4ce11b47d328b3babca76a734a65c171f5

Observation c3e1d71b-d402-466c-a341-5baf90e30292 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mastering the game of go with deep neural networks and tree search

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.840989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.840989Z digest=sha256:b67a591c476930e6b97b290c1525843f22a2d5400ee789ce2e70e1375730bbb0

Observation a01d5683-b46c-40b4-9454-fdbea02541dc · outbound

This paper cites Bayes' Bluff: Opponent Modelling in Poker.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Bayes' Bluff: Opponent Modelling in Poker

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.017512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.017512Z digest=sha256:0f2fa2f174d0257711ebb1a4ade231b4d59574d2658c875841acfffbfd45c587

Observation 2eb77aa2-f1af-4a39-8bb1-bc87abb3f955 · outbound

This paper cites Brown, Adam Santoro, Aditya Gupta, et al.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Brown, Adam Santoro, Aditya Gupta, et al

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.402090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:26.131722Z digest=sha256:543ce415292641d99b41407effcde375d16398103e7555698b25776346df000c

Observation 8c0b0d61-6407-4193-b951-8fee35869ffe · outbound

This paper cites Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.207323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.207323Z digest=sha256:2ba0ec21c0f4d5b55a564fc16897a68fd233d6f95b9826184722ca79f470c963

Observation ad54ce16-f332-43d9-82dd-ddde84dc1c39 · outbound

This paper cites Le, Ed H.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Le, Ed H

Reference 29

Resolution
malformed identifier
no resolver link, observed 2026-08-05T14:23:26.320509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.320509Z digest=sha256:7097b3b0e45dfad00e10978b6b47523c1bd922c76a40f1b4a084400af6267a92

Observation 7cb96572-77c7-4a03-bca6-b78143d72d34 · outbound

This paper cites CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.402477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.402477Z digest=sha256:c613b888f7e9310eb559459a14b6029d93c5c38c9c4f9a0517b621d7b97ddfa9

Observation e1d6f15b-32c1-4415-95b7-cce9431820eb · outbound

This paper cites StarCraft II: A New Challenge for Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models StarCraft II: A New Challenge for Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.486476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.486476Z digest=sha256:5c58aad1a8e48ceeb5e9879b13dcb2acfb75ddde96601e53a2c9802f4d390304

Observation 6ec25b2f-cfd9-4b6c-a1f6-3fa2eeb994c4 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.587398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.587398Z digest=sha256:80fd8a4615b5d81ea55d058a2d31835e1ba2b939a87a962c17f20c8ab44afc6b

Observation 6ef93714-881f-4b35-80d7-bf99c4f62fd2 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.680923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.680923Z digest=sha256:46678bd2956f9c6241d6b1506acf6efcd51ad88a7045ede6bde473aa95a7ac27

Observation 1f538ca5-df5c-4728-8f1e-c26e3ae4beff · outbound

This paper cites Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.747722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.747722Z digest=sha256:b561e1c5bfcd5098f2aa21367bd81e252d9b0e716edbf2ac16868d348e9fb885

Observation e9e3842a-54be-4d53-beb0-b66881e83778 · outbound

This paper cites Agents Play Thousands of 3D Video Games.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Agents Play Thousands of 3D Video Games

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.797595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.797595Z digest=sha256:447edfc16170b902fae8da70c3fa1aa37eedfdfc61418d2512bc6c26116185d0

Observation 5dbef2c2-1c83-443d-93f3-8b5e6f3c6e76 · outbound

This paper cites Policy-to-language: Train llms to explain decisions with flow-matching generated rewards.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Policy-to-language: Train llms to explain decisions with flow-matching generated rewards

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-05T14:23:28.595131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:26.982264Z digest=sha256:d294561c65471db42486e6bbef915ef09c37b42c0e25894f5daeb7dbc23df421

Observation e6705ad7-3bbe-4ea4-a259-aaef503742f2 · outbound

This paper cites Mastering complex control in moba games with deep reinforcement learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mastering complex control in moba games with deep reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.251013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:27.199232Z digest=sha256:ae0a46babd93023a414205fd2eac83727431de525a6bc53f20d6db9b1b9b5a66

Observation 34de5102-2fc4-47ed-a428-202845f6c721 · outbound

This paper cites C har P oet: A C hinese classical poetry generation system based on token-free LLM.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models C har P oet: A C hinese classical poetry generation system based on token-free LLM

Reference 38

Resolution
verified exact
doi, observed 2026-08-05T14:23:27.806325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:27.276998Z digest=sha256:6448ca3889daa9f7cdb4020321258930f1089c5ebcc9168f2a6f11b233a56e76

Observation 5a7d02ff-d198-40fd-88b0-c455de45c09f · outbound

This paper cites Training interactive agent in large fps game map with rule-enhanced reinforcement learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Training interactive agent in large fps game map with rule-enhanced reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.097095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:23:27.367053Z digest=sha256:023a487e2c1709e61585109726e1db628bf0dadea1146d3a6fc9539f20cbf41a

Observation 0aeb653d-cc59-49f5-ac7c-a1e24532cf66 · outbound

This paper cites Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:27.444486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:27.444486Z digest=sha256:1a8c1083477f0516d12974470306056273f1613d8c0e13633c5419691ab3b266

Observation 4007ec55-0812-434b-9b84-7b42004e4305 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:27.530464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:27.530464Z digest=sha256:bfb97a892719e3e6631616f710515762ce39ed2165424e09455269cb9c06c340

Observation 9f49b6f6-5ae9-4788-8807-a56fc5849d5e · outbound

This paper cites PokerBench: Training Large Language Models to become Professional Poker Players.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models PokerBench: Training Large Language Models to become Professional Poker Players

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:27.591640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:27.591640Z digest=sha256:77bd1d1882a13e115e1fe18128ee29e2662ab0595a61dc52fe6e6543c0492e69

Pith citing papers

Observation 682ba848-7a4b-4e77-a5d3-b9c69669f7e6 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

Reference 288

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:50:48.309200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:ef24bb3614ac7ca3a9565261c1e13c5751bf364512ac6a65f862681b2885dcd3

Observation f1e7d76b-4a07-474e-8ef8-18cc3018c509 · inbound

SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction cites this paper.

SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:01.106317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:34:01.106317Z digest=sha256:54f8d915fd951a3f32dda3bea5706985e48dc90af9f64b0294aa123ce4b1b860