Pith. sign in

REVIEW 8 cited by

It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.18570 v3 pith:KG5PKRB6 submitted 2024-05-28 cs.CV cs.CLcs.IRcs.LG

classification cs.CVcs.CLcs.IRcs.LG
keywords contrastivespacecliplossmodalityembeddingsmulti-modalrepresentational
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modal contrastive models such as CLIP achieve state-of-the-art performance in zero-shot classification by embedding input images and texts on a joint representational space. Recently, a modality gap has been reported in two-encoder contrastive models like CLIP, meaning that the image and text embeddings reside in disjoint areas of the latent space. Previous studies suggest that this gap exists due to 1) the cone effect, 2) mismatched pairs in the dataset, and 3) insufficient training. We show that, even when accounting for all these factors, and even when using the same modality, the contrastive loss actually creates a gap during training. As a result, We propose that the modality gap is inherent to the two-encoder contrastive loss and rename it the contrastive gap. We present evidence that attributes this contrastive gap to low uniformity in CLIP space, resulting in embeddings that occupy only a small portion of the latent space. To close the gap, we adapt the uniformity and alignment properties of unimodal contrastive loss to the multi-modal setting and show that simply adding these terms to the CLIP loss distributes the embeddings more uniformly in the representational space, closing the gap. In our experiments, we show that the modified representational space achieves better performance than default CLIP loss in downstream tasks such as zero-shot image classification and multi-modal arithmetic.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

    cs.CL 2026-05 conditional novelty 7.0 of 10

    TokenSwap measures and mitigates the MLLM modality gap: swapping textual concepts for matched images lowers accuracy by 4-47% across 42 models, and training with such swaps reduces the gap.

  2. On the modality gap and the contrastive loss in multi-modal representation learning

    cs.LG 2026-07 conditional novelty 6.5 of 10

    InfoNCE with independent encoders actively creates a modality gap at low temperature; mixing intra- and inter-modality negatives (xNCE) removes the gap while improving zero-shot transfer.

  3. Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MG-CLIP preserves CLIP's modality gap by adaptively limiting fine-tuning epochs and compensates for its limits with a visual-space classifier, improving class-incremental learning without replay.

  4. PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning

    cs.MM 2025-07 conditional novelty 6.0 of 10

    Keeping only the first 12 layers of Qwen2-VL plus self-distillation and a modality-aware contrastive loss yields a 3B unified multimodal retriever within 1.8 points of the 7B model on M-BEIR.

  5. Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

    cs.IR 2025-06 conditional novelty 6.0 of 10

    Vela adapts an audio MLLM into a universal text-audio embedding model using 'in one word' prompts, in-context examples, and text-only contrastive training, outperforming CLAP-style models on retrieval benchmarks.

  6. On the rankability of visual embeddings

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Visual embeddings from CLIP and other vision encoders encode ordinal attributes along linear directions, recoverable from as few as two extreme reference images, without full supervision.

  7. Aligning Multimodal Representations through an Information Bottleneck

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A regularizer derived from an information-bottleneck bound, essentially a mean-squared alignment loss, reduces modality-specific information and improves multimodal alignment and image captioning.

  8. Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A background-profile subtraction method improves zero-shot sound classification accuracy under noisy conditions, and narrowing the audio-text modality gap further boosts performance.

Pith tools