A 7B vision-language model trained with reinforcement learning in a new four-game environment, VLM-Gym, outperforms larger proprietary models and improves its perception and reasoning abilities together.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning
A 7B vision-language model trained with reinforcement learning in a new four-game environment, VLM-Gym, outperforms larger proprietary models and improves its perception and reasoning abilities together.