Pith. sign in

Paper Citation Record · LEDGER

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 12 inbound Pith citation observations for arXiv:2505.13426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13426 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:40.641794Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:12.064728Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 632309ab-ceac-4221-83a5-87b08d498007 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.340425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.414846Z digest=sha256:4ff78efbad285c056762c306b59b020a92034a15106e387963e616595083eeec

Observation 24134682-60cd-461a-b811-80fa3ec1cec0 · outbound

This paper cites Qwen2.5-VL Technical Report.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.420400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.420400Z digest=sha256:c001ed4a01490b8b4ac8e4c40779e97279c3c7ce260a78e0198b7e5febb232bc

Observation 279d5243-cd23-4ba9-affe-6dd70ebc6c66 · outbound

This paper cites Deep blue.Artificial intelligence, 134(1-2):57–83, 2002.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Deep blue.Artificial intelligence, 134(1-2):57–83, 2002

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.425329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.425329Z digest=sha256:86ca388ca521428c974277c34269e28d445d5f59ea1847df5a87ba2338367128

Observation 3fec197c-b599-4d52-8449-4d313ca7646f · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.313219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.430240Z digest=sha256:4a800f9b2e8282de7f9a50976e4b381b952430e0e52955d09c4ab0a6eff602d1

Observation f83a8618-2e4b-4f9d-80ce-38bf088742d0 · outbound

This paper cites Next token prediction towards multimodal intelligence: A comprehensive survey, 2024.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Next token prediction towards multimodal intelligence: A comprehensive survey, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.297446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.435063Z digest=sha256:9a0dbe92a59a3f31c024b0d4f795d1c25fd6b7d42cae45c7831f0015267a6bef

Observation 369135b3-c9ed-42dd-8a5e-39620b1cf007 · outbound

This paper cites Pca-bench: Evaluating multimodal large language models in perception-cognition-action chain, 2024.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Pca-bench: Evaluating multimodal large language models in perception-cognition-action chain, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.281126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.439988Z digest=sha256:cbf3669d6ac9d8b2054377330c13f59c04cb67eec72b3a9f8fc8785989968f26

Observation 1a6d9cb4-f391-46e4-aaeb-b0cca2fb1a16 · outbound

This paper cites Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.444733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.444733Z digest=sha256:06d7cc9d4cab1a62f13eb445ce47c1a1ed76be87c32b3145ac89e89a1af90838

Observation efd27942-730f-4235-bcc4-794e94b13ce8 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.264761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.449752Z digest=sha256:c651a8a6e61d2af1e92a63e016feb598b19311f0f9e175959da12d8c6c4e2115

Observation 67ab1055-ff75-4e04-977b-dcc0f3ffde41 · outbound

This paper cites GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.454784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.454784Z digest=sha256:374e7dc17e8813a591b3418088beee70a20f0688ccc8530714bc418e256dc36a

Observation 1c4d1528-a51b-4f91-89a1-44542b429875 · outbound

This paper cites FlowReasoner: Reinforcing Query-Level Meta-Agents.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning FlowReasoner: Reinforcing Query-Level Meta-Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.459605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.459605Z digest=sha256:c85cc9b6bc151040b24845dce0a8ea98eda229d51e4d2ab4d642e9f7a9a25675

Observation c351b643-633e-4b3b-8341-6134f6b7f1a8 · outbound

This paper cites VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.464359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.464359Z digest=sha256:2b7dc8d6aacd474319345745d4fb0584c32d7a6649e1a669fff4002d72b20e8f

Observation 9669f4b8-7798-426c-a7bb-20595bb0425e · outbound

This paper cites Mmevalpro: Calibrating multimodal benchmarks towards trustworthy and efficient evaluation, 2025.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Mmevalpro: Calibrating multimodal benchmarks towards trustworthy and efficient evaluation, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.249782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.469372Z digest=sha256:78b5332f7385c7aef793d9910b70d2f0b394bec4eb627f0eddae86ba9f3ac824

Observation 8f7321ce-28c0-413c-8135-2ff7b7a81489 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.473965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.473965Z digest=sha256:7e26c9fa8815bee7501d824dc512a44096d148390d0c9068b8925b907e945a43

Observation 4dc716f5-6070-476c-85a1-73e1b3daf1c9 · outbound

This paper cites Learning multiple layers of features from tiny images.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Learning multiple layers of features from tiny images

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.478709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.478709Z digest=sha256:cb8add1204f8d44900091d9cf92be15b307bff02857a9859b063e16ba6cb7bda

Observation 5854a8e8-56d2-44d4-8836-6ee06d98abf4 · outbound

This paper cites JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.483549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.483549Z digest=sha256:f467bf41d487347ebd20d3ee35d0df5f13ebd4bff8531af6eabec88afd6f7d97

Observation 2acb931c-f258-4e73-a7b0-304dbf0627d0 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.488243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.488243Z digest=sha256:2a855f27141eaf0a7a34eeaff2876342e4857d424a24c1594f1a0729137997de

Observation 554de052-0550-4e9e-8ce0-6b395012918b · outbound

This paper cites Playing atari with deep reinforcement learning, 2013.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Playing atari with deep reinforcement learning, 2013

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.492543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.492543Z digest=sha256:a580bcbd69584a9055a317dc93008c85cb407779c27720fea00d48a77f1422e6

Observation 0520e4b8-2b95-49be-bc5d-68c3f11b8e40 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.496919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.496919Z digest=sha256:487e2bed0403cfed23ea391b4388c47506ab217cb954c0a42ba62bcaadd22410

Observation 1a33699d-6603-444a-bd60-73df5d3a53f2 · outbound

This paper cites Gpt-4v(ision) system card.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Gpt-4v(ision) system card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.501709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.501709Z digest=sha256:c6bb2cc3632e14633ec9f0871aba87b7333bed03862d721aa059dff1168f241e

Observation fa0cbb04-9ab0-4d0b-8e6c-9f2ed2d4b218 · outbound

This paper cites hello-gpt-4o, 2024.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning hello-gpt-4o, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.197712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.506425Z digest=sha256:5bcc6faca98cbef74d9532fc83520411fcb4b5e6b2d147d3a4b8cbbd7e357295

Observation 4bb05e01-a1b2-4078-8281-07f3f2c907c3 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.510952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.510952Z digest=sha256:2f0164c2f9ba2388cfbca9fc7813f03f0538065bba0f30965133d823a1276a53

Observation 937eff5f-c397-41e7-9755-fe62a921f8a5 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Generative agents: Interactive simulacra of human behavior

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.182835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.516003Z digest=sha256:ea210d52cbf1053a9ba28575f340d17efec6053340114e52a6b096c844ac86cc

Observation 8360e5ef-6d8b-4245-9f09-f53d9aac24f2 · outbound

This paper cites Llms are greedy agents: Effects of rl fine-tuning on decision-making abilities, 2025.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Llms are greedy agents: Effects of rl fine-tuning on decision-making abilities, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.520512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.520512Z digest=sha256:9446fd4c14c215c92be8b5f7262b51b9793fdb09c00007370c14cfdbb527dec3

Observation feeea2f0-153e-4c9d-9859-c3dff08eeac0 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.525076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.525076Z digest=sha256:392c887e5e4c2d35b6bd596081b2f9a1d2dd4c76933982f3f75ddb5e67abde71

Observation aed021ba-1089-49ad-b927-6d7003827db8 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.529437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.529437Z digest=sha256:e37feb8bd955204a62e88da4fb48fe070c62c8e635179c778a1c72104feda256

Observation 37e6be08-a1fd-43a2-b726-7067a8694a7d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.533997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.533997Z digest=sha256:563d84d0b8e77a3042e58b047cbe983791925d7c85f81ff1067337dca1372c95

Observation e0eaf1a4-9c0b-49a7-88a5-79cf97e50fc1 · outbound

This paper cites Maddison, Arthur Guez, L.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Maddison, Arthur Guez, L

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.538357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.538357Z digest=sha256:70b7e37105f6d246fbcabc5e8f90ddaffd07779f1c3cb105710fb3f970d6423b

Observation 5a6cad25-9eab-41cf-82e2-5771c660ffbb · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.543185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.543185Z digest=sha256:3f8292379d8c333be695102740b965c61173734292cbfa557c5b303d5c8da2aa

Observation e768e2fc-e5e3-4485-bdd2-97e99ee52e0d · outbound

This paper cites Cradle: Empowering Foundation Agents Towards General Computer Control.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Cradle: Empowering Foundation Agents Towards General Computer Control

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.547708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.547708Z digest=sha256:acbf3e4f22d878eeeaced714497880aec24859f5420cf90be47b19dbe91a02b1

Observation 75fc430f-ef66-4d7d-bd53-b8427f8ad369 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.122308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.552327Z digest=sha256:ed06f20ab9116e9be5cf14c12e5294b24f61b07f095ec5fec0f567d502ce157b

Observation 4082641c-36e1-4276-ada1-bc9f938798ce · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.556940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.556940Z digest=sha256:88637bb3b1dd321110e572bc52ecd43c5e4ad8e60162ce5833dd7e24eaf05268

Observation 33a03cbd-d4af-4138-b4a4-741a8feb7f18 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.561840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.561840Z digest=sha256:2881e96b85c72056f658f28294ae34a34a191b512177567683f60d36788ce9d1

Observation 5c4bc011-e5b3-4f25-bb36-774bc6d6d5e7 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.097633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.566256Z digest=sha256:871a187fab0adca293c4fcd1ddf3fcbe289a6c8fa4110005916945d18ff892ac

Observation 31b151e0-0773-48b5-b0e1-ef04023fb839 · outbound

This paper cites Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.570782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.570782Z digest=sha256:a3ff10cfbc5b9016d9467978a56573974c07b6094a25be0292e1e481cbf4bc5e

Observation 5389e299-c36a-421e-b639-b6871230a2fa · outbound

This paper cites Are large vision language models good game players?, 2025.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Are large vision language models good game players?, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.575106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.575106Z digest=sha256:45b2a605b4b2ed45b9e1d94716365499c98226aeb49b465b55cc5700848f76d2

Observation 7a376b86-59df-4aeb-b987-b5ce65fecef5 · outbound

This paper cites Are Large Vision Language Models Good Game Players?.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Are Large Vision Language Models Good Game Players?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.579494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.579494Z digest=sha256:955bca95aeffe1fa5bf0aff36ce2978e063ed309424621a0e9e1b99e7bde60b2

Observation e8c5e5f7-f051-476a-84ad-7f46718e2760 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.584123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.584123Z digest=sha256:1eabafc920490f67c5e3063a48ae385873a5f35b1ec523673cc0c406b09e9bc4

Observation 88cd7c0a-e4be-42db-8f9b-9298ab9bb032 · outbound

This paper cites Waytowich, Devin White, MD Sunbeam, and Vinicius G.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Waytowich, Devin White, MD Sunbeam, and Vinicius G

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.073150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.588424Z digest=sha256:351dd2b6b60317a3e3efb5a30dd392095e6af4ca30293b3fd94bb8ef461108ee

Observation 3c5628f0-b510-41b6-afb7-d5133c637cf5 · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.592511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.592511Z digest=sha256:4d4803754ead6c4e1e0c9eee43cb2f75b7258bd8ccab4d46ba9b06af2a6ae0cb

Observation 49c54ada-c4b2-4a80-bd78-2547dfcf2463 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.596690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.596690Z digest=sha256:a0e43c729058465f72e49777504e028ebae42791ae6730c58f041f277bdc82ce

Observation 4f129bf9-27fd-4024-8774-5aec1655a0ac · outbound

This paper cites Fine-tuning large vision-language models as decision-making agents via reinforcement learning.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Fine-tuning large vision-language models as decision-making agents via reinforcement learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.059318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.601099Z digest=sha256:c8fb75abbfde578282c8cb3e2e8dbba524c8d30f8b10ff852656dd6439cb4585

Observation 30b7371c-8f7b-4be9-ba6d-6c5884c4719a · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.044872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.605736Z digest=sha256:d7b8b1b7fd5f55f42ac9ec94f4b9fee5e239f4bd6f62582ec0a01c6d64af88dd

Observation 16eff4c8-22ad-4de0-b4b0-9a1c1f59bb4d · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.610285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.610285Z digest=sha256:5f2de51b2ec98dc76af02ffb826577ff07fa9edf6b0cf3082fdc76aa2373798a

Observation 192e4191-6ac2-43c1-8814-1fbb634907e1 · outbound

This paper cites aha moment.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning aha moment

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:41.019518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.615135Z digest=sha256:fde7cafa19a6a564b18f88438c5f1653522377dca9517881906bbda8142f53df

Observation 742fffcb-b53a-4d78-9b03-812468680c52 · outbound

This paper cites Output Format Description First describe the board in <perception></perception>.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Output Format Description First describe the board in <perception></perception>

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:40.986156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.624234Z digest=sha256:9cf813991db11774420250cbb67e9922b6e87c20bb27494c775a2658a59786b5

Observation 5946c486-f3bb-4bcd-8867-09cd084c5fa0 · outbound

This paper cites (row1,col1) (row2,col2).

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning (row1,col1) (row2,col2)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:40.970513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.632844Z digest=sha256:0141db9cbb496fcf8fc28ce013750194f3aff2e4843e3606794664b3726dcfdf

Observation af7af115-aad6-4ed2-9647-c2359f32d881 · outbound

This paper cites an unresolved cited work.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:41.001590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.637255Z digest=sha256:a36a3ae2f819b419db2e43a608c636d8a61370616c3b00d3e84a32702bfc99e3

Observation ce5ad0eb-10e1-4d55-855b-a909c6ce8a03 · outbound

This paper cites Output Format Description First describe the board in <perception></perception>.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning Output Format Description First describe the board in <perception></perception>

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:40.955184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:21:40.641794Z digest=sha256:31cfa18f4c9ec1d3eb3b60533673a2da5a2a0be740f83c6ded9e2596ec8e47c0

Pith citing papers

Observation b6a2d442-edd3-43fe-b415-f99bd643a4e8 · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.064728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.064728Z digest=sha256:f206ad07eb02ca13b15161f928028bf810aa1a8934e28b13c8fdf0d57d41431f

Observation 562becc3-0bf1-4cd9-b8f0-829ea6af43a6 · inbound

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning cites this paper.

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:21:53.290875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T22:17:41.758059Z digest=sha256:09cf3597a8d50672777355f224a5bd57556a64093e736b2f0ac63efb275f9bfd

Observation be41000b-fa61-4774-89c7-5dbac2f99424 · inbound

Explain Before You Answer: A Survey on Compositional Visual Reasoning cites this paper.

Explain Before You Answer: A Survey on Compositional Visual Reasoning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 159

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:18.102569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:18.102569Z digest=sha256:75320ac70b8ae99b012da27f68aa5ee4e8bf52493cdbd87b722462d649a3d735

Observation 3c920247-9f73-4a74-b31d-de700b27b369 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.400095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:951ba120fcba2db07b0924200d6d8b0830d59b32e5b4dff81eb33e8f940e2b3d

Observation bd874107-7871-46f7-a6e0-00b32ff204d2 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.071498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:ac2485d363135f87a07d33931193c3c4442321bd0f351cd09270091b97771000

Observation 956c631a-0ef5-4bf5-8f4f-2c7720d85ba0 · inbound

Gym-V: A Unified Vision Environment System for Agentic Vision Research cites this paper.

Gym-V: A Unified Vision Environment System for Agentic Vision Research G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:05:26.028670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T10:05:09.049846Z digest=sha256:eee2caa4df3e5fa40f6c61560d4c03998140d33e031813f14d81a7e148641939

Observation 58c72bb5-6fd8-4fc1-87b9-46224ea591f5 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:08.788471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:723db6e58e8268b181ea9386a1e00cf4737b221251e5f75b6ce982bf3f4a3142

Observation c91ea36f-f7fd-4ea2-bbdc-c0320bde6f33 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.110488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:178f2616b1d970b1543ae3402ded161c8127e5786444f22dc255bcfc29d8de6e

Observation 9da3c042-ac19-4f9d-b2fa-fe0fe1f26042 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.994016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:711a5e244988d843d25d11f9219f1b3e07924708f16949eacaf4ae622a078b4c

Observation 69838789-0ba6-4fc1-aaae-0b6f5097178e · inbound

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning cites this paper.

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:13:05.251364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T04:09:32.397341Z digest=sha256:49874a8b4c0e65cedaf89eb6425e3ab2959145c58e49edb7a84352f1d2fc6a53

Observation 3f6d3dd1-ceba-448d-8587-55ad11df59ff · inbound

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning cites this paper.

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:41.170593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T05:55:19.083517Z digest=sha256:666825e73e10244441c100ab3beb4b188ec125dd21efc43f3af2b956d0855dae

Observation 2e6a6b4c-2bcb-419c-8434-262108d00612 · inbound

CAST: Game Solvers as Turn-Level Teachers for LLM Agents cites this paper.

CAST: Game Solvers as Turn-Level Teachers for LLM Agents G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:58:29.293555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:58:29.293555Z digest=sha256:cd8c48c7463c62febd978eb8fed1d442e82fcb7fc62ee7566cfe880bcb8351cf