Pith. sign in

REVIEW 4 cited by

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.05464 v2 pith:R6SZF4T4 submitted 2025-05-08 cs.CL

classification cs.CL
keywords mergingreasoningmodelsperceptionmodellayersabilitiescapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-Language Models (VLMs) combine visual perception with the general capabilities, such as reasoning, of Large Language Models (LLMs). However, the mechanisms by which these two abilities can be combined and contribute remain poorly understood. In this work, we explore to compose perception and reasoning through model merging that connects parameters of different models. Unlike previous works that often focus on merging models of the same kind, we propose merging models across modalities, enabling the incorporation of the reasoning capabilities of LLMs into VLMs. Through extensive experiments, we demonstrate that model merging offers a successful pathway to transfer reasoning abilities from LLMs to VLMs in a training-free manner. Moreover, we utilize the merged models to understand the internal mechanism of perception and reasoning and how merging affects it. We find that perception capabilities are predominantly encoded in the early layers of the model, whereas reasoning is largely facilitated by the middle-to-late layers. After merging, we observe that all layers begin to contribute to reasoning, whereas the distribution of perception abilities across layers remains largely unchanged. These observations shed light on the potential of model merging as a tool for multimodal integration and interpretation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Recursive Vision Language Models for General Symbolic Reasoning

    cs.CV 2026-08 conditional novelty 5.0 of 10

    R-Qwen, a LoRA-adapted Qwen model that iteratively refines explicit candidate solutions under constraint projection, outperforms prior recursive models and zero-shot frontier LLMs on eight symbolic reasoning benchmarks.

  2. VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    Merging a coding LLM into a vision-language model via task vectors yields an open-source multimodal coder that reaches near-GPT-4o performance on the authors' new benchmark.

  3. Training-Free Reasoning and Reflection in MLLMs

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Training-free, layer-wise weight merging of an MLLM with a reasoning LLM, using attention-derived priors, raises MMMU accuracy from 63.9 to 69.2 at the 38B scale.

  4. PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models

    cs.CL 2025-05 reject novelty 5.0 of 10

    A domain-specialized 7B vision-language model improves most planning-map interpretation scores on a new 300-example benchmark, but the headline 'best overall' claim is contradicted by the same table.

Pith tools