Pith. sign in

REVIEW 3 cited by

Unveiling the Tapestry of Consistency in Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14156 v4 pith:7MEBMC4Q submitted 2024-05-23 cs.CV

classification cs.CV
keywords consistencylvlmsmodelssolutionanswersconbenchaccuracycaption
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information. However, when faced with prompts in different sizes of solution spaces, LVLMs fail to always give consistent answers regarding the same knowledge point. This inconsistency of answers between different solution spaces is prevalent in LVLMs and erodes trust. To this end, we provide a multi-modal benchmark ConBench, to intuitively analyze how LVLMs perform when the solution space of a prompt revolves around a knowledge point. Based on the ConBench tool, we are the first to reveal the tapestry and get the following findings: (1) In the discriminate realm, the larger the solution space of the prompt, the lower the accuracy of the answers. (2) Establish the relationship between the discriminative and generative realms: the accuracy of the discriminative question type exhibits a strong positive correlation with its Consistency with the caption. (3) Compared to open-source models, closed-source models exhibit a pronounced bias advantage in terms of Consistency. Eventually, we ameliorate the consistency of LVLMs by trigger-based diagnostic refinement, indirectly improving the performance of their caption. We hope this paper will accelerate the research community in better evaluating their models and encourage future advancements in the consistency domain. The project is available at https://github.com/foundation-multimodal-models/ConBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AbilityLens unifies 11 public benchmarks into six perception abilities with accuracy and stability metrics, and reveals ability conflicts during MLLM training linked to data mixing and model size.

  2. On the Consistency of Video Large Language Models in Temporal Comprehension

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Most open-source Video-LLMs are near chance-level at verifying their own temporal predictions, and a proposed event-temporal verification tuning improves both grounding and consistency.

  3. Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.

Pith tools