Pith. sign in

Paper Citation Record · LEDGER

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 12 inbound Pith citation observations for arXiv:2505.13426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13426 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:40.641794Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:12.064728Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 632309ab-ceac-4221-83a5-87b08d498007 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.340425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.414846Z digest=sha256:34ada3a48f14f7476c9725bcb85acdffce1f21b3750a1c36a75f22bbed5de97d

Observation 24134682-60cd-461a-b811-80fa3ec1cec0 · outbound

This paper cites Qwen2.5-VL Technical Report.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.420400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.420400Z digest=sha256:eca2227268d76e6d94a755865dced9bbdc9f535364c89bd8009c8d24bda5e2c3

Observation 279d5243-cd23-4ba9-affe-6dd70ebc6c66 · outbound

This paper cites Deep blue.Artificial intelligence, 134(1-2):57–83, 2002.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Deep blue.Artificial intelligence, 134(1-2):57–83, 2002

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.425329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.425329Z digest=sha256:adab610225de8d5a4ddf49543a2197edc5f449a0b1ba5fadec909a5f5b667253

Observation 3fec197c-b599-4d52-8449-4d313ca7646f · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.313219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.430240Z digest=sha256:c3f76945eedf8f5bd8cac783423f6349874fa74e9dca8bee6bfa889a142c1589

Observation f83a8618-2e4b-4f9d-80ce-38bf088742d0 · outbound

This paper cites Next token prediction towards multimodal intelligence: A comprehensive survey, 2024.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Next token prediction towards multimodal intelligence: A comprehensive survey, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.297446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.435063Z digest=sha256:4cb52337dd31f9f53ff7e2b83beaa920b49c3f557f2a79a4a7fd2bb10f785e63

Observation 369135b3-c9ed-42dd-8a5e-39620b1cf007 · outbound

This paper cites Pca-bench: Evaluating multimodal large language models in perception-cognition-action chain, 2024.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Pca-bench: Evaluating multimodal large language models in perception-cognition-action chain, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.281126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.439988Z digest=sha256:ed1a7e202d14057e687146ae998fecf024a9699e0419cb289cfdd99236fabfcd

Observation 1a6d9cb4-f391-46e4-aaeb-b0cca2fb1a16 · outbound

This paper cites Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.444733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.444733Z digest=sha256:45966b0df371a2d5ec5a0876427fce0999d6144fa5e84495e494e067588a7781

Observation efd27942-730f-4235-bcc4-794e94b13ce8 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.264761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.449752Z digest=sha256:09bb6b70c845d48460bc4cb049dd9bc68aeddb2fdae032602c7cf1bcb10fed9f

Observation 67ab1055-ff75-4e04-977b-dcc0f3ffde41 · outbound

This paper cites GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.454784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.454784Z digest=sha256:f49f8ce5095c1e132ec0d73647296fad155a41db2febae82f7f23fa237582262

Observation 1c4d1528-a51b-4f91-89a1-44542b429875 · outbound

This paper cites FlowReasoner: Reinforcing Query-Level Meta-Agents.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning FlowReasoner: Reinforcing Query-Level Meta-Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.459605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.459605Z digest=sha256:1cdf82f8c9e752cf60d2a66274694d120414cf273261c6387bdbf78f5d891765

Observation c351b643-633e-4b3b-8341-6134f6b7f1a8 · outbound

This paper cites VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.464359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.464359Z digest=sha256:b6b1c904db1b7c64691fff32d20595be12f8420f138c955a51efae826fcf212f

Observation 9669f4b8-7798-426c-a7bb-20595bb0425e · outbound

This paper cites Mmevalpro: Calibrating multimodal benchmarks towards trustworthy and efficient evaluation, 2025.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Mmevalpro: Calibrating multimodal benchmarks towards trustworthy and efficient evaluation, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.249782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.469372Z digest=sha256:212ccc4b8edfb127d17df5c6b690c5be73767bd3f88c76cf8ad7a2233fc273e8

Observation 8f7321ce-28c0-413c-8135-2ff7b7a81489 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.473965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.473965Z digest=sha256:31c13536ad9e65984d54c964424bb9022af8a19a1d3337f35db26793be2d0319

Observation 4dc716f5-6070-476c-85a1-73e1b3daf1c9 · outbound

This paper cites Learning multiple layers of features from tiny images.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Learning multiple layers of features from tiny images

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.478709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.478709Z digest=sha256:acb311cf563206390977ee3d146fb38017383ace1f7e8fce6a85ef0c5bb5dfaa

Observation 5854a8e8-56d2-44d4-8836-6ee06d98abf4 · outbound

This paper cites JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.483549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.483549Z digest=sha256:4cafcc34d3168a24ad710cc187e45eccf210fab3307ec850728e18f7b6106b88

Observation 2acb931c-f258-4e73-a7b0-304dbf0627d0 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.488243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.488243Z digest=sha256:ffa99747f28e368c794a8d2aa7f7891a4a045c4461639ae3fad4b3c8536b1c05

Observation 554de052-0550-4e9e-8ce0-6b395012918b · outbound

This paper cites Playing atari with deep reinforcement learning, 2013.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Playing atari with deep reinforcement learning, 2013

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.492543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.492543Z digest=sha256:ab05f1c36d7870d84e0daa98f759ef51e72abfcd11e3e2dd3927a18053ea4e7e

Observation 0520e4b8-2b95-49be-bc5d-68c3f11b8e40 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.496919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.496919Z digest=sha256:c5b281ac4ee925ad0f36d9e9cd56ebc0a9d6ae313dc5f66b18d5c57ee0602860

Observation 1a33699d-6603-444a-bd60-73df5d3a53f2 · outbound

This paper cites Gpt-4v(ision) system card.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Gpt-4v(ision) system card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.501709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.501709Z digest=sha256:03d674e1964256a59b6d50fbe5816423e1fb8e623f32f1414ab7ee0e75d10aeb

Observation fa0cbb04-9ab0-4d0b-8e6c-9f2ed2d4b218 · outbound

This paper cites hello-gpt-4o, 2024.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning hello-gpt-4o, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.197712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.506425Z digest=sha256:345880589527c4dc1aaf7ea6c5cb4b57b34bb4135e839f324633d54f287b2f0c

Observation 4bb05e01-a1b2-4078-8281-07f3f2c907c3 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.510952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.510952Z digest=sha256:dbf01be592711aef2c896e6fb4fabe6c0abf4c8bb1d5f314afbc5ef6cc8f15a0

Observation 937eff5f-c397-41e7-9755-fe62a921f8a5 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Generative agents: Interactive simulacra of human behavior

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.182835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.516003Z digest=sha256:d9fb19c477ad88779b2669590850d22647b21dcab1ed33d9ab30206d6bee59f1

Observation 8360e5ef-6d8b-4245-9f09-f53d9aac24f2 · outbound

This paper cites Llms are greedy agents: Effects of rl fine-tuning on decision-making abilities, 2025.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Llms are greedy agents: Effects of rl fine-tuning on decision-making abilities, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.520512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.520512Z digest=sha256:f1c242a2145b64fcaa81d649ff8598bce8b16835b3a2b1cf769f78fad7c3ba17

Observation feeea2f0-153e-4c9d-9859-c3dff08eeac0 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.525076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.525076Z digest=sha256:c86414a1ae405a87b0cc6e34189fab6e07e77c59c70bc2172b7d2e13ff32168d

Observation aed021ba-1089-49ad-b927-6d7003827db8 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.529437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.529437Z digest=sha256:a6189ff897172f57356d21befd1d8d741ad572a4ea581a2f9de91dabaceaef81

Observation 37e6be08-a1fd-43a2-b726-7067a8694a7d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.533997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.533997Z digest=sha256:68f444a5601d416a81b4a1c8874e60449a99cbd208a6ccf46db760d0424358fc

Observation e0eaf1a4-9c0b-49a7-88a5-79cf97e50fc1 · outbound

This paper cites Maddison, Arthur Guez, L.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Maddison, Arthur Guez, L

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.538357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.538357Z digest=sha256:627198b5879757c28079483d674dbb7c363232ad26c124ff12e5404b07455d48

Observation 5a6cad25-9eab-41cf-82e2-5771c660ffbb · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.543185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.543185Z digest=sha256:9cae84510fbfb0aee5c7c178492b6c29360f266c1f4609b13ae7b9a7aa2b4102

Observation e768e2fc-e5e3-4485-bdd2-97e99ee52e0d · outbound

This paper cites Cradle: Empowering Foundation Agents Towards General Computer Control.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Cradle: Empowering Foundation Agents Towards General Computer Control

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.547708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.547708Z digest=sha256:496148ec6ba8f265c0b9758ea1446353b721fa54b51445761b801279b469de85

Observation 75fc430f-ef66-4d7d-bd53-b8427f8ad369 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.122308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.552327Z digest=sha256:c4ec180639d8cc8b6c22dbbe00d783ac3356cce16744198f1faeebe502a9ad71

Observation 4082641c-36e1-4276-ada1-bc9f938798ce · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.556940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.556940Z digest=sha256:1bdc1f83016b0351ddc7641470972a042f2e14d1eaf05ad569915a4aa277bb19

Observation 33a03cbd-d4af-4138-b4a4-741a8feb7f18 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.561840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.561840Z digest=sha256:eab3bf672f78f9e8626a61df6816a3263092b28780a96ebd9f8c1f4719bb44c1

Observation 5c4bc011-e5b3-4f25-bb36-774bc6d6d5e7 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.097633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.566256Z digest=sha256:a89145622a884179586709429a39d5aa2836461c19cef8b6821bb8f2823b1123

Observation 31b151e0-0773-48b5-b0e1-ef04023fb839 · outbound

This paper cites Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.570782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.570782Z digest=sha256:f6245be241f97d2a05b9ebca62569a6d056a19d546e7e475c24550a1b3d02fa2

Observation 5389e299-c36a-421e-b639-b6871230a2fa · outbound

This paper cites Are large vision language models good game players?, 2025.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Are large vision language models good game players?, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.575106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.575106Z digest=sha256:54c03782311bc1ac19be15604cf9b4890ab1186b3e5d0200920b4ff6cfe11130

Observation 7a376b86-59df-4aeb-b987-b5ce65fecef5 · outbound

This paper cites Are Large Vision Language Models Good Game Players?.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Are Large Vision Language Models Good Game Players?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.579494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.579494Z digest=sha256:52f8d238d7622369bfdec5c68568f48e7605fcef03743c690ce5fe512cf94b47

Observation e8c5e5f7-f051-476a-84ad-7f46718e2760 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.584123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.584123Z digest=sha256:231a8f662eec78452b63bbb4e61b6978319efbe13acaf96e2be66b18f2ab1318

Observation 88cd7c0a-e4be-42db-8f9b-9298ab9bb032 · outbound

This paper cites Waytowich, Devin White, MD Sunbeam, and Vinicius G.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Waytowich, Devin White, MD Sunbeam, and Vinicius G

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.073150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.588424Z digest=sha256:b7a8ceb918e837d2d1780ef620dfce73ab8dbcca93e307366b92001980c34948

Observation 3c5628f0-b510-41b6-afb7-d5133c637cf5 · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.592511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.592511Z digest=sha256:046ebc1a59bca1d7049b30206972bdab0cac61bbd072bb18cea898a3cfe5ec4b

Observation 49c54ada-c4b2-4a80-bd78-2547dfcf2463 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.596690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.596690Z digest=sha256:787921cb914fe2c1ac40e473c4dc45b20fc03e29941eeaecb9164e0ce06f5c80

Observation 4f129bf9-27fd-4024-8774-5aec1655a0ac · outbound

This paper cites Fine-tuning large vision-language models as decision-making agents via reinforcement learning.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Fine-tuning large vision-language models as decision-making agents via reinforcement learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.059318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.601099Z digest=sha256:aed4cb38d798bfeed1fc029ca24c94baee4a03d6f9db87a669e87d81d09f1207

Observation 30b7371c-8f7b-4be9-ba6d-6c5884c4719a · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.044872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.605736Z digest=sha256:a9d00db449f1829cbd638e58f05eb796eec4107e2067bb9b7618318fd8aebf86

Observation 16eff4c8-22ad-4de0-b4b0-9a1c1f59bb4d · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.610285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.610285Z digest=sha256:302ac81536a7a7bbbb062d076e66245bf35f603b5ac60c5fac060d8209c91b25

Observation 192e4191-6ac2-43c1-8814-1fbb634907e1 · outbound

This paper cites aha moment.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning aha moment

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.019518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.615135Z digest=sha256:64a2d583aa9a86f4da90c9dc654f7f27de17953773d31fe1e55f8049428377c7

Observation 742fffcb-b53a-4d78-9b03-812468680c52 · outbound

This paper cites Output Format Description First describe the board in <perception></perception>.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Output Format Description First describe the board in <perception></perception>

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:40.986156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.624234Z digest=sha256:a1d85d0f399ba77dc76e92ce80371c1a24cc68bb24810df8267b0f13754d15e2

Observation 5946c486-f3bb-4bcd-8867-09cd084c5fa0 · outbound

This paper cites (row1,col1) (row2,col2).

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning (row1,col1) (row2,col2)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:40.970513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.632844Z digest=sha256:6740ab277398f28de349c4cd56e7c3b505fdd599d53d506222d1818729649314

Observation af7af115-aad6-4ed2-9647-c2359f32d881 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.001590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.637255Z digest=sha256:54422694e8f60ac16a898fd69923fd192e242930c93475a4f87177a70237b8a2

Observation ce5ad0eb-10e1-4d55-855b-a909c6ce8a03 · outbound

This paper cites Output Format Description First describe the board in <perception></perception>.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Output Format Description First describe the board in <perception></perception>

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:40.955184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:40.641794Z digest=sha256:e9569920bdc1a41591ed1bd6463638259d3eb205b24d4f39bae5cdb29e336dfd

Pith citing papers

Observation b6a2d442-edd3-43fe-b415-f99bd643a4e8 · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.064728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.064728Z digest=sha256:2dfe2453b74a84ea56b957e956780513bb5b23a8733765da40f0bf4b69140da7

Observation 562becc3-0bf1-4cd9-b8f0-829ea6af43a6 · inbound

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning cites this paper.

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:21:53.290875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T22:17:41.758059Z digest=sha256:1c684a735aba97116e6e7da59dbced863c990f7ca3db71ad3091cd9b65baec01

Observation be41000b-fa61-4774-89c7-5dbac2f99424 · inbound

Explain Before You Answer: A Survey on Compositional Visual Reasoning cites this paper.

Explain Before You Answer: A Survey on Compositional Visual Reasoning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 159

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:18.102569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:18.102569Z digest=sha256:55b14c2565f215b0b29e27d40e49dab6cb41a2f5db84dd4f441cfd7bc4c68975

Observation 3c920247-9f73-4a74-b31d-de700b27b369 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.400095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:d5813fd69402447ad172a912ff53b57208493275b0f11c6e19bd0a965011185f

Observation bd874107-7871-46f7-a6e0-00b32ff204d2 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.071498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:5a7357cd0e33fbf8b5c27bd42a772e223e9e01b307fe0129cf47be574000f235

Observation 956c631a-0ef5-4bf5-8f4f-2c7720d85ba0 · inbound

Gym-V: A Unified Vision Environment System for Agentic Vision Research cites this paper.

Gym-V: A Unified Vision Environment System for Agentic Vision Research G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:05:26.028670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T10:05:09.049846Z digest=sha256:fff2ddb812948db272f0b20fff87bd1039b43ea7b0b06e1cbd23a0fe6b5caf23

Observation 58c72bb5-6fd8-4fc1-87b9-46224ea591f5 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:08.788471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:6345dbe50fedef7e77e6eddedca01eead6dc7f157aedc0a6a1d586f9ee1f0c12

Observation c91ea36f-f7fd-4ea2-bbdc-c0320bde6f33 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.110488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:c6c05607bfa83dc10ca12e77bbfb8e4ec6319fe627fc17688128032082be2d04

Observation 9da3c042-ac19-4f9d-b2fa-fe0fe1f26042 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.994016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:7799bdd29d2340dee6cd06a416930ade945a2fb59ea52ad828b313b7aed362c6

Observation 69838789-0ba6-4fc1-aaae-0b6f5097178e · inbound

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning cites this paper.

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:13:05.251364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T04:09:32.397341Z digest=sha256:60afa000d5db27851d03e1009960a15af1ec126342362edd57800cdfc2e6e650

Observation 3f6d3dd1-ceba-448d-8587-55ad11df59ff · inbound

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning cites this paper.

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:41.170593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T05:55:19.083517Z digest=sha256:044d9ddcf4b872c2017765c6f781d6cd838257e6d791965bf89470a1c74335e8

Observation 2e6a6b4c-2bcb-419c-8434-262108d00612 · inbound

CAST: Game Solvers as Turn-Level Teachers for LLM Agents cites this paper.

CAST: Game Solvers as Turn-Level Teachers for LLM Agents G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:58:29.293555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:58:29.293555Z digest=sha256:5075205f97aaf86ef38b123d18beaa22fde830718c90556df2bfa71a885b495b