Pith. sign in

REVIEW 4 cited by

BRAVE: Broadening the visual encoding of vision-language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.07204 v1 pith:CI3LRQZG submitted 2024-04-10 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords visualvlmsdifferentencodersbiasesbraveencodingfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language models (VLMs) are typically composed of a vision encoder, e.g. CLIP, and a language model (LM) that interprets the encoded features to solve downstream tasks. Despite remarkable progress, VLMs are subject to several shortcomings due to the limited capabilities of vision encoders, e.g. "blindness" to certain image features, visual hallucination, etc. To address these issues, we study broadening the visual encoding capabilities of VLMs. We first comprehensively benchmark several vision encoders with different inductive biases for solving VLM tasks. We observe that there is no single encoding configuration that consistently achieves top performance across different tasks, and encoders with different biases can perform surprisingly similarly. Motivated by this, we introduce a method, named BRAVE, that consolidates features from multiple frozen encoders into a more versatile representation that can be directly fed as the input to a frozen LM. BRAVE achieves state-of-the-art performance on a broad range of captioning and VQA benchmarks and significantly reduces the aforementioned issues of VLMs, while requiring a smaller number of trainable parameters than existing methods and having a more compressed representation. Our results highlight the potential of incorporating different visual biases for a more broad and contextualized visual understanding of VLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Exam for Active Observers

    cs.CV 2026-07 conditional novelty 7.0 of 10

    On a new 17-task benchmark of active visual observation, the best frontier multimodal model solves 10.6% of items and humans solve 96.1%.

  2. Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Stable Diffusion features, especially when conditioned on the question, improve vision-centric multimodal question answering when fused with CLIP.

  3. NanoVLMs: How small can we go and still make coherent Vision Language Models?

    cs.CV 2025-02 reject novelty 5.0 of 10

    NanoVLMs, 5M to 25M parameter vision-language models trained on simplified GPT-4o captions, are judged by GPT-4o as nearly as coherent as the 50x larger Kosmos-2 on a 25-sample test.

  4. Multi-Agent Interactive Question Generation Framework for Long Document Understanding

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A multi-agent question generation pipeline produces long-context English and Arabic QA pairs (AraEngLongBench), and top LVLMs score below 50% on the resulting benchmark.

Pith tools