Pith. sign in

REVIEW 7 cited by

Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.07983 v3 pith:SBBRV7SI submitted 2024-04-11 cs.CV cs.LG

classification cs.CVcs.LG
keywords biasobjectmodalityothercontrastivefindinformationlike
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks. Despite their successes in some tasks, like zero-shot object recognition, they perform surprisingly poor on other tasks, like attribute recognition. Previous work has attributed these challenges to the modality gap, a separation of image and text in the shared representation space, and to a bias towards objects over other factors, such as attributes. In this analysis paper, we investigate both phenomena thoroughly. We evaluated off-the-shelf VLMs and while the gap's influence on performance is typically overshadowed by other factors, we find indications that closing the gap indeed leads to improvements. Moreover, we find that, contrary to intuition, only few embedding dimensions drive the gap and that the embedding spaces are differently organized. To allow for a clean study of object bias, we introduce a definition and a corresponding measure of it. Equipped with this tool, we find that object bias does not lead to worse performance on other concepts, such as attributes per se. However, why do both phenomena, modality gap and object bias, emerge in the first place? To answer this fundamental question and uncover some of the inner workings of contrastive VLMs, we conducted experiments that allowed us to control the amount of shared information between the modalities. These experiments revealed that the driving factor behind both the modality gap and the object bias, is an information imbalance between images and captions, and unveiled an intriguing connection between the modality gap and entropy of the logits.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

    cs.CL 2026-05 conditional novelty 7.0 of 10

    TokenSwap measures and mitigates the MLLM modality gap: swapping textual concepts for matched images lowers accuracy by 4-47% across 42 models, and training with such swaps reduces the gap.

  2. Defect-aware Hybrid Prompt Optimization via Progressive Tuning for Zero-Shot Multi-type Anomaly Detection and Segmentation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    DAPO learns shared defect-aware text prompts that let CLIP segment known and novel defect types in unseen industrial domains.

  3. Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

    cs.CV 2025-02 conditional novelty 6.0 of 10

    CLIP's intra-modal embeddings are miscalibrated: converting one side of an image-image or text-text comparison into the other modality improves retrieval accuracy on 15+ datasets.

  4. AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization

    cs.CV 2026-07 reject novelty 5.0 of 10

    AspectCLIP partitions captions into four SimCSE-based clusters and restricts cyclic consistency regularization to within clusters, reporting modest gains over CyCLIP on several CLIP benchmarks.

  5. When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

    cs.CV 2026-02 conditional novelty 5.0 of 10

    VLAs fail most counterfactual instructions because vision shortcuts dominate language; the new LIBERO-CF benchmark quantifies this, and CAG, an inference-time action mixer, improves grounding and success.

  6. Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.

  7. Aligning Multimodal Representations through an Information Bottleneck

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A regularizer derived from an information-bottleneck bound, essentially a mean-squared alignment loss, reduces modality-specific information and improves multimodal alignment and image captioning.

Pith tools