Pith. sign in

REVIEW 4 cited by

A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12436 v2 pith:O6DMJXIC submitted 2023-12-19 cs.CV cs.AIcs.CLcs.MM

classification cs.CVcs.AIcs.CLcs.MM
keywords geminigpt-4vvisualcapabilitiesmllmmulti-modalchallengerearly
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow Large Language Models (LLMs) with powerful capabilities in visual understanding, enabling them to tackle diverse multi-modal tasks. Very recently, Google released Gemini, its newest and most capable MLLM built from the ground up for multi-modality. In light of the superior reasoning capabilities, can Gemini challenge GPT-4V's leading position in multi-modal learning? In this paper, we present a preliminary exploration of Gemini Pro's visual understanding proficiency, which comprehensively covers four domains: fundamental perception, advanced cognition, challenging vision tasks, and various expert capacities. We compare Gemini Pro with the state-of-the-art GPT-4V to evaluate its upper limits, along with the latest open-sourced MLLM, Sphinx, which reveals the gap between manual efforts and black-box systems. The qualitative samples indicate that, while GPT-4V and Gemini showcase different answering styles and preferences, they can exhibit comparable visual reasoning capabilities, and Sphinx still trails behind them concerning domain generalizability. Specifically, GPT-4V tends to elaborate detailed explanations and intermediate steps, and Gemini prefers to output a direct and concise answer. The quantitative evaluation on the popular MME benchmark also demonstrates the potential of Gemini to be a strong challenger to GPT-4V. Our early investigation of Gemini also observes some common issues of MLLMs, indicating that there still remains a considerable distance towards artificial general intelligence. Our project for tracking the progress of MLLM is released at https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis

    cs.CV 2025-05 reject novelty 6.0 of 10

    Internal attention patterns in multimodal LLMs are used to define an attention accuracy metric and a benchmark for detecting cases where a model answers correctly while attending to the wrong image.

  2. RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A rationale-aware defensive prompting framework uses multimodal chain-of-thought and self-checking to reduce harmful MLLM outputs while preserving benign utility.

  3. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

    cs.CV 2024-12 conditional novelty 5.0 of 10

    The authors integrate streaming perception, compressed long-term memory, and a reasoning model into one open-source system, reporting SOTA open-source results on several video and audio benchmarks.

  4. Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey

    cs.CL 2024-11 conditional novelty 1.0 of 10

    A survey that organizes VQA methods from feature extraction through MLLM reasoning, datasets, and metrics, without introducing new experimental results.

Pith tools