Pith. sign in

REVIEW 9 cited by

Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03659 v2 pith:CGJTKNHI submitted 2024-10-04 cs.CV cs.CL

classification cs.CVcs.CL
keywords conflictsknowledgemodelscontrastiveaccuracycomponentsconflictcross-modality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities for capturing and reasoning over multimodal inputs. However, these models are prone to parametric knowledge conflicts, which arise from inconsistencies of represented knowledge between their vision and language components. In this paper, we formally define the problem of $\textbf{cross-modality parametric knowledge conflict}$ and present a systematic approach to detect, interpret, and mitigate them. We introduce a pipeline that identifies conflicts between visual and textual answers, showing a persistently high conflict rate across modalities in recent LVLMs regardless of the model size. We further investigate how these conflicts interfere with the inference process and propose a contrastive metric to discern the conflicting samples from the others. Building on these insights, we develop a novel dynamic contrastive decoding method that removes undesirable logits inferred from the less confident modality components based on answer confidence. For models that do not provide logits, we also introduce two prompt-based strategies to mitigate the conflicts. Our methods achieve promising improvements in accuracy on both the ViQuAE and InfoSeek datasets. Specifically, using LLaVA-34B, our proposed dynamic contrastive decoding improves an average accuracy of 2.24%.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A new taxonomy and 1,500-item dataset, ENTRAP-VL, lets researchers measure whether vision-language models are entrained by textual and visual context separately.

  2. Linguistic Context Recodes Visual Representations in Vision-Language Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Goal-directed language prompts make VLMs add a transferable goal-relevant marker to selected image objects and amplify those objects' queried attributes in later layers, and both effects causally influence answers.

  3. How Do Vision-Language Models Process Conflicting Information Across Modalities?

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Vision-language models answer from whichever modality is encoded more saliently in their final-layer representations, and specific attention heads can be manipulated to shift that preference.

  4. Is Extending Modality The Right Path Towards Omni-Modality?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Fine-tuning LLMs on extra modalities improves some knowledge tasks but degrades reasoning and instruction-following; weighted model merging preserves language ability better than training one model on all modalities.

  5. mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A systematic empirical study finds that for multimodal RAG, EVA-CLIP retrieval, listwise LVLM reranking, and feeding only the top-ranked document works best, with a self-reflection agent adding further gains.

  6. Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MMKC-Bench provides a human-verified benchmark of multimodal knowledge conflicts and shows that current LMMs prefer internal parametric knowledge over external evidence.

  7. Robust Multimodal Large Language Models Against Modality Conflict

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A new benchmark, MMMC, triggers hallucinations in all tested multimodal models, and reinforcement learning on it reduces such hallucinations more than prompt engineering or supervised fine-tuning.

  8. MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models

    cs.CV 2025-07 conditional novelty 4.0 of 10

    MCA-LLaVA reindexes image tokens by sums of mirrored 2D coordinates so instruction tokens attend across the whole image, reducing hallucination on POPE, CHAIR, and MME.

  9. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0 of 10

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...

Pith tools