Pith. sign in

REVIEW 8 cited by

Vision language models are blind: Failing to translate detailed visual features into words

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.06581 v6 pith:BG6KOY6C submitted 2024-07-09 cs.AI cs.CV

classification cs.AIcs.CV
keywords modelsvisionvlmsaccuracyinformationlanguagetasksblindtest
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While large language models with vision capabilities (VLMs), e.g., GPT-4o and Gemini 1.5 Pro, score high on many vision-understanding benchmarks, they are still struggling with low-level vision tasks that are easy to humans. Specifically, on BlindTest, our suite of 7 very simple tasks, including identifying (a) whether two circles overlap; (b) how many times two lines intersect; (c) which letter is being circled in a word; and (d) the number of circles in an Olympic-like logo, four state-of-the-art VLMs are only 58.07% accurate on average. Claude 3.5 Sonnet performs the best at 77.84% accuracy, far from the human expected accuracy of 100%. Across different image resolutions and line widths, VLMs including slow-thinking models consistently struggle with those tasks that require precise spatial information when geometric primitives overlap or are close. Yet, VLMs perform at near-100% accuracy when much more space is added to separate shapes and letters. Linear probing experiments show that vision encoders contain sufficient visual information to solve BlindTest and that language models fail to decode this information into correct answers. Code and data are at: https://vlmsareblind.github.io

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Exam for Active Observers

    cs.CV 2026-07 conditional novelty 7.0 of 10

    On a new 17-task benchmark of active visual observation, the best frontier multimodal model solves 10.6% of items and humans solve 96.1%.

  2. The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    Task conditioning suppresses safety-critical signal reporting in language and vision models that unconstrained versions report at higher rates, creating an inattentional gap that decouples benchmark safety from real-w...

  3. Vision Language Models Cannot Reason About Physical Transformation

    cs.AI 2026-03 accept novelty 6.5 of 10

    Current VLMs cannot maintain transformation-invariant representations of number, length, volume or size and instead rely on textual invariance priors that reverse on matched non-conserving controls.

  4. Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.

  5. Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 23-dimension benchmark finds that even the best VLMs score near random on motion trajectory, temporal extension, and several prediction tasks, far below humans, suggesting weak internal world models.

  6. Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test

    cs.AI 2025-05 conditional novelty 6.0 of 10

    With chain-of-thought prompting and text inputs, GPT-4o, Gemini-1.5 Pro, and Claude-3.5 Sonnet reach or exceed human-level set-shifting on the WCST, but not with visual inputs or direct answers.

  7. Teach Me Sign: Stepwise Prompting LLM for Sign Language Production

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Fine-tuning an LLM with GPT-4o-generated sign language structure assistance improves sign pose generation over a Progressive Transformer baseline on Phoenix14T and How2Sign.

  8. MiMo-VL Technical Report

    cs.CL 2025-06 conditional novelty 5.0 of 10

    MiMo-VL-7B-RL, a 7B open-source vision-language model, reports state-of-the-art results on 35 of 40 benchmarks and a 59.4 OlympiadBench score, with the report crediting long-CoT pretraining data and mixed on-policy RL.

Pith tools