Pith. sign in

Paper Citation Record · LEDGER

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2508.21365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21365 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:27.591640Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:01.106317Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact6
  • verified fuzzy6
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 3462f19f-6220-49a1-a503-049fa24058ce · outbound

This paper cites Cause and Effect: Can Large Language Models Truly Understand Causality?.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Cause and Effect: Can Large Language Models Truly Understand Causality?

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:23:28.247417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:23.843562Z digest=sha256:75ecbfa15abd492bdf34e2db1b540b72a9814efe70236e80b2c2ebd502494d23

Observation 6039e9bf-ca8c-441f-97b9-581a2cd2d618 · outbound

This paper cites an unresolved cited work.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:23.929458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:23.929458Z digest=sha256:32adec8ae669031ba98666639a9b327400e0d178d6e611927bfd4e113a81a5ef

Observation 9b2c6e68-646a-43be-92f7-971cf3f8ff64 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.020818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.020818Z digest=sha256:bfb840bf67abee99c8bd615e1d29b996939dcf8dbf2973bf15b1cc7b9bdb3f91

Observation 65eeae8c-36e2-48ce-8939-8ceb51a810c7 · outbound

This paper cites MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.072935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.072935Z digest=sha256:b35eec6b4768c1c2181a1fca4411b25550cb8012f28f8b6ccacfaa9ffa0bda5c

Observation 7a428987-6583-4222-9603-fb9898fd35bc · outbound

This paper cites Bayeschess: A computer chess program based on bayesian networks.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Bayeschess: A computer chess program based on bayesian networks

Reference 5

Resolution
verified exact
doi, observed 2026-08-05T14:23:28.080495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.156489Z digest=sha256:8b8777e402e80ea3ddfbf342b3e3785122502fe93ddd2bb20b56276f14582012

Observation d4b144ae-8637-4860-883c-01ca553fd71a · outbound

This paper cites Font and Tobias Mahlmann.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Font and Tobias Mahlmann

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T14:23:29.926244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.207525Z digest=sha256:968a0d7c411669b81295c645cdbb84d645b5ba2a1056f1bf29253360d80bc590

Observation ce16c2fa-817c-4f36-9a6e-3bc437a6a1b5 · outbound

This paper cites Enabling self-improving agents to learn at test time with human-in-the-loop guidance.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Enabling self-improving agents to learn at test time with human-in-the-loop guidance

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-05T14:23:29.665817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.271531Z digest=sha256:ea3413f57b4060b55170a5abd693ba7574d4a0f2aa75baca04f56f2477e181a2

Observation 7bb921e6-e04c-48b4-83d7-f04eef122fb0 · outbound

This paper cites Measuring massive multitask language understanding.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Measuring massive multitask language understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.347757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.347757Z digest=sha256:148dd8b8460ef0f58ebdbd9cd28d05e3e29d6fd8766b7d7529c570856e94c1f0

Observation 712a2b16-a84f-48c1-95db-a3c78fcec02a · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.434666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.434666Z digest=sha256:4bf1f4fd5dea10c645e86b2d603387c74a5bb1faeb3d9eb54d5edc3de004928e

Observation 4d97786b-a523-4a9f-a8cb-623b7b38872f · outbound

This paper cites A Survey on Large Language Model-Based Game Agents.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models A Survey on Large Language Model-Based Game Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.473717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.473717Z digest=sha256:541189ab72e334efd942d0287cd8595053d6b75528f6f3f4fe51e8002f130d34

Observation 37e15f8e-9e20-42f3-b515-e38236208f3f · outbound

This paper cites PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.541597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.541597Z digest=sha256:ba4ccf60bdd074fba8d67f6a0909b6a154e33c426bdd5a20342495b5d5d44179

Observation b120537f-e788-4299-b64b-8bc3f74f543e · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.920211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.602432Z digest=sha256:16de17d5b222738f3e322067bb3072f48c1ab83f282ab5f4c8275ebbc4c12d21

Observation 0a47fa39-e61c-4bd8-90a1-6ee078b5988d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.667850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.667850Z digest=sha256:a7a69b6c6e9195bdba2a1e05da998df3033813edf6a224902f6d9d0916ab1e35

Observation b8d24579-c921-478d-a8ca-613a937f29ae · outbound

This paper cites Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:23:29.386562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.726570Z digest=sha256:fd1c127fbba65a8695982e0cf98756badcbeb61a36665d8a4b16a37be6ba0e55

Observation a7266d34-80bb-40ec-a956-b4b069ea27d6 · outbound

This paper cites School chinese benchmark, 2018.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models School chinese benchmark, 2018

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.727816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.798984Z digest=sha256:220e2c9eda2bdcfd5f8dc1ae2837dbffb4bb678a63038a86553309b60c2a22fb

Observation 4d0d153b-c670-4065-85c6-183d48962969 · outbound

This paper cites CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.911042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.911042Z digest=sha256:c621b95ea9c9e124ea04c23b6d679ace4393894e8190d7082d605390e077460d

Observation d0e6f647-06ac-422b-be93-43d35f0c9b2e · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Simpo: Simple preference optimization with a reference-free reward

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.975898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.975898Z digest=sha256:0a619942b92d40376f36a3ecee3540e52372b427ce150fae6b44961d369882db

Observation 8fe91552-3333-4fac-b23d-823e54708096 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Playing Atari with Deep Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.071210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.071210Z digest=sha256:a3bac5e3c783198393e6dc350f7a2f40f88e061bd9cc2b59cbc6bb2b426b19fb

Observation 2ad2ca66-1586-4690-931c-d94db8641ab5 · outbound

This paper cites Creating pro-level AI for a real-time fighting game using deep reinforcement learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Creating pro-level AI for a real-time fighting game using deep reinforcement learning

Reference 19

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T14:23:29.160079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:25.134829Z digest=sha256:01ce2358e3009f19d94ba4c5f357027e696cdb4accc52a10ae77a9299f19332a

Observation 30ed4bd9-ea79-4473-95b0-6d637aab7769 · outbound

This paper cites Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, et al.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.589098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:25.311280Z digest=sha256:10243deef75f65ca62515415d5f363af3e2c8a2b3130e6d872135609cc53a955

Observation cf1e4c3f-0971-4329-965f-2e2f5a9a7df6 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Manning, Stefano Ermon, and Chelsea Finn

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.379171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.379171Z digest=sha256:8c346dd69e9e306f942808c9edaf6f7de3c7ac91f3713ba449826b9b0622f352

Observation 95e7b3ed-3fcd-47db-89ce-a6f1e284b427 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.463390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.463390Z digest=sha256:70c8f8eeba4cd9c8e4c6d1b911fc7de9450bac993a786307c94d2febcf2f1b34

Observation 6ad11b99-f8f3-4a63-b322-a6221a23776a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.587387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.587387Z digest=sha256:9c5d45f6dbdb93fe3dc7878ccbb3c1282d4d0126d8e2da407d99e63eadbec523

Observation 5678c3e7-8d6f-494f-b41c-ad52a791efca · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.667622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.667622Z digest=sha256:da224c8c81cae5adfdff8663f2858187786617ff3be903ecef2bbaa7c90798b7

Observation c3e1d71b-d402-466c-a341-5baf90e30292 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mastering the game of go with deep neural networks and tree search

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.840989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.840989Z digest=sha256:979096f048bea888528fb8934d723c98f753fcd72268420078c44da94ef5e74a

Observation a01d5683-b46c-40b4-9454-fdbea02541dc · outbound

This paper cites Bayes' Bluff: Opponent Modelling in Poker.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Bayes' Bluff: Opponent Modelling in Poker

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.017512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.017512Z digest=sha256:10dc5e22cf8d65a57cd0be8d2e83294939ae2f128facb146ea3afe78e2610fc3

Observation 2eb77aa2-f1af-4a39-8bb1-bc87abb3f955 · outbound

This paper cites Brown, Adam Santoro, Aditya Gupta, et al.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Brown, Adam Santoro, Aditya Gupta, et al

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.402090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:26.131722Z digest=sha256:1916f6027e55c9293494578cc368d55070f7c3168e41f5f6c06444a842743e7f

Observation 8c0b0d61-6407-4193-b951-8fee35869ffe · outbound

This paper cites Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.207323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.207323Z digest=sha256:228c184347575d34fcfc801bfe324e4910e32e82746134ee73e039310af268ba

Observation ad54ce16-f332-43d9-82dd-ddde84dc1c39 · outbound

This paper cites Le, Ed H.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Le, Ed H

Reference 29

Resolution
malformed identifier
no resolver link, observed 2026-08-05T14:23:26.320509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.320509Z digest=sha256:4f91848fa5d32389e1ddf4220e98fc0bc6e9daf0612dfb47e74f8142235a7e1e

Observation 7cb96572-77c7-4a03-bca6-b78143d72d34 · outbound

This paper cites CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.402477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.402477Z digest=sha256:4b94cdeac3cd111e992aa9d262e59c07511332e67aad3fe6364e3e12151f43bb

Observation e1d6f15b-32c1-4415-95b7-cce9431820eb · outbound

This paper cites StarCraft II: A New Challenge for Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models StarCraft II: A New Challenge for Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.486476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.486476Z digest=sha256:18c4ed878a2516b113fb72aa5d1ef4f9a8f864f25f5234f0b7da518a56a3fa1b

Observation 6ec25b2f-cfd9-4b6c-a1f6-3fa2eeb994c4 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.587398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.587398Z digest=sha256:b19115e77494a58cee8ef563542610d2445ad02f377f6ccc675c02de4f154d5c

Observation 6ef93714-881f-4b35-80d7-bf99c4f62fd2 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.680923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.680923Z digest=sha256:66170230c992dfe24d8309c0db95578b353f78505fb6e8c55c0643c5da61ea79

Observation 1f538ca5-df5c-4728-8f1e-c26e3ae4beff · outbound

This paper cites Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.747722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.747722Z digest=sha256:b8d73d054a277d3ded6a830b18e8c7e09c7aef1f9128390863d0b6210c239114

Observation e9e3842a-54be-4d53-beb0-b66881e83778 · outbound

This paper cites Agents Play Thousands of 3D Video Games.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Agents Play Thousands of 3D Video Games

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.797595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.797595Z digest=sha256:6defdcb5747d36f5163e4118716eaf6a7f9b75f55f6cd1c49d58a60bd3a2d06a

Observation 5dbef2c2-1c83-443d-93f3-8b5e6f3c6e76 · outbound

This paper cites Policy-to-language: Train llms to explain decisions with flow-matching generated rewards.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Policy-to-language: Train llms to explain decisions with flow-matching generated rewards

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-05T14:23:28.595131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:26.982264Z digest=sha256:131d36e8b7baa32c777f41ac0b12e4d8a9d6296eadcd81c6142395035318aa19

Observation e6705ad7-3bbe-4ea4-a259-aaef503742f2 · outbound

This paper cites Mastering complex control in moba games with deep reinforcement learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mastering complex control in moba games with deep reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.251013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:27.199232Z digest=sha256:22798c995b7a1f83c339e6485922442f2d16f22b1dadd123519aaffb1acdbbbc

Observation 34de5102-2fc4-47ed-a428-202845f6c721 · outbound

This paper cites C har P oet: A C hinese classical poetry generation system based on token-free LLM.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models C har P oet: A C hinese classical poetry generation system based on token-free LLM

Reference 38

Resolution
verified exact
doi, observed 2026-08-05T14:23:27.806325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:27.276998Z digest=sha256:975eb71bc581baf24cb08a4d753aba13553db7fb48ef63986dec12828b303e10

Observation 5a7d02ff-d198-40fd-88b0-c455de45c09f · outbound

This paper cites Training interactive agent in large fps game map with rule-enhanced reinforcement learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Training interactive agent in large fps game map with rule-enhanced reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.097095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T14:23:27.367053Z digest=sha256:aa8b0f88e6fd403cfe5f2cc4ba8af4cdf56470b69d176b5e7058c9d092a04797

Observation 0aeb653d-cc59-49f5-ac7c-a1e24532cf66 · outbound

This paper cites Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:27.444486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:27.444486Z digest=sha256:66a2e798a9ce9cdb0901bd607701fce5fdbda4798f536257bc9c3dda1b914356

Observation 4007ec55-0812-434b-9b84-7b42004e4305 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:27.530464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:27.530464Z digest=sha256:5755ba89c755030362ed0b777ca2850f40e0bbeee2e89563b73ea88f308528b9

Observation 9f49b6f6-5ae9-4788-8807-a56fc5849d5e · outbound

This paper cites PokerBench: Training Large Language Models to become Professional Poker Players.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models PokerBench: Training Large Language Models to become Professional Poker Players

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:27.591640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:27.591640Z digest=sha256:525a2283d0bb9202e16dc996411461225945a69af7813027486fdfcaba00d654

Pith citing papers

Observation 682ba848-7a4b-4e77-a5d3-b9c69669f7e6 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

Reference 288

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:50:48.309200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:02b25cfc26b036bf1f9131d910423030861c3ec50ea68f5974c26df9ae81a3a1

Observation f1e7d76b-4a07-474e-8ef8-18cc3018c509 · inbound

SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction cites this paper.

SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:01.106317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:34:01.106317Z digest=sha256:a6345f013235131eae8a53239aa5d9c3be183d176efc9ed9c34ad880b4bac183