Pith. sign in

REVIEW 1 cited by

A Surprising Failure? Multimodal LLMs and the NLVR Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17793 v1 pith:RJ4XG3LT submitted 2024-02-26 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords nlvrcompositionalimagemodelreasoningsentencetaskbiases
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This study evaluates three state-of-the-art MLLMs -- GPT-4V, Gemini Pro, and the open-source model IDEFICS -- on the compositional natural language vision reasoning task NLVR. Given a human-written sentence paired with a synthetic image, this task requires the model to determine the truth value of the sentence with respect to the image. Despite the strong performance demonstrated by these models, we observe they perform poorly on NLVR, which was constructed to require compositional and spatial reasoning, and to be robust for semantic and systematic biases.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning

    cs.AI 2025-06 conditional novelty 6.0 of 10

    State-of-the-art multimodal language models perform at or near random chance on MARBLE, a new hard benchmark for spatial reasoning and planning.

Pith tools