Training VLMs to explicitly convert images to text before reasoning transfers simple-to-hard generalization from text to image, and this conversion skill can be internalized to keep inference cheap.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
Training VLMs to explicitly convert images to text before reasoning transfers simple-to-hard generalization from text to image, and this conversion skill can be internalized to keep inference cheap.