Pith. sign in

REVIEW 2 cited by

InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01487 v1 pith:CEYM26CG submitted 2024-03-03 cs.CV

classification cs.CV
keywords infimm-hdmllmshigh-resolutionimagesmodelsmultimodalvisualaccurate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal Large Language Models (MLLMs) have experienced significant advancements recently. Nevertheless, challenges persist in the accurate recognition and comprehension of intricate details within high-resolution images. Despite being indispensable for the development of robust MLLMs, this area remains underinvestigated. To tackle this challenge, our work introduces InfiMM-HD, a novel architecture specifically designed for processing images of different resolutions with low computational overhead. This innovation facilitates the enlargement of MLLMs to higher-resolution capabilities. InfiMM-HD incorporates a cross-attention module and visual windows to reduce computation costs. By integrating this architectural design with a four-stage training pipeline, our model attains improved visual perception efficiently and cost-effectively. Empirical study underscores the robustness and effectiveness of InfiMM-HD, opening new avenues for exploration in related areas. Codes and models can be found at https://huggingface.co/Infi-MM/infimm-hd

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A two-stage crop-and-predict framework improves high-resolution MLLM performance by using the model's own coarse localization to focus on a candidate region before final prediction.

  2. Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A 3B VLM trained with GRPO to call a zoom tool improves V*Bench accuracy by 5.7% over its base model but degrades TextVQA and HR-Bench performance.

Pith tools