REVIEW 6 cited by
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper explores the problem of commonsense level vision-knowledge conflict in Multimodal Large Language Models (MLLMs), where visual information contradicts model's internal commonsense knowledge. To study this issue, we introduce an automated framework, augmented with human-in-the-loop quality control, to generate inputs designed to simulate and evaluate these conflicts in MLLMs. Using this framework, we have crafted a diagnostic benchmark consisting of 374 original images and 1,122 high-quality question-answer (QA) pairs. The benchmark covers two aspects of conflict and three question types, providing a thorough assessment tool. We apply this benchmark to assess the conflict-resolution capabilities of nine representative MLLMs from various model families. Our results indicate an evident over-reliance on parametric knowledge for approximately 20% of all queries, especially among Yes-No and action-related problems. Based on these findings, we evaluate the effectiveness of existing approaches to mitigating the conflicts and compare them to our "Focus-on-Vision" prompting strategy. Despite some improvement, the vision-knowledge conflict remains unresolved and can be further scaled through our data construction framework. Our proposed framework, benchmark, and analysis contribute to the understanding and mitigation of vision-knowledge conflicts in MLLMs.
Forward citations
Cited by 6 Pith papers
-
Prior Bias in Vision Language Models on UML Diagram Interpretation
Reversing only the UML relation arrow while keeping class names and layout fixed cuts open-source VLM relation accuracy by about 33%, revealing prior-over-vision bias.
-
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
A new taxonomy and 1,500-item dataset, ENTRAP-VL, lets researchers measure whether vision-language models are entrained by textual and visual context separately.
-
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
MMKC-Bench provides a human-verified benchmark of multimodal knowledge conflicts and shows that current LMMs prefer internal parametric knowledge over external evidence.
-
MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts
Ten leading VLMs mostly fail to report removed essential object parts as missing, and simulated detector evidence, image tools, longer reasoning, and an easier fine-tune barely improve accuracy.
-
Robust Multimodal Large Language Models Against Modality Conflict
A new benchmark, MMMC, triggers hallucinations in all tested multimodal models, and reinforcement learning on it reduces such hallucinations more than prompt engineering or supervised fine-tuning.
-
MLLMs are Deeply Affected by Modality Bias
A position paper with a case study showing that multimodal LLMs rely on language priors and underuse visual input, together with a research roadmap and calls for balanced training.
Discussion (0). Continue with ORCID to comment.