Pith. sign in

REVIEW 1 cited by

Multi-Modal Representation Learning with Text-Driven Soft Masks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.00719 v1 pith:FKOYVUCJ submitted 2023-04-03 cs.CV

classification cs.CV
keywords learningmulti-modaldatadiverseexamplesframeworkimage-textloss
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a visual-linguistic representation learning approach within a self-supervised learning framework by introducing a new operation, loss, and data augmentation strategy. First, we generate diverse features for the image-text matching (ITM) task via soft-masking the regions in an image, which are most relevant to a certain word in the corresponding caption, instead of completely removing them. Since our framework relies only on image-caption pairs with no fine-grained annotations, we identify the relevant regions to each word by computing the word-conditional visual attention using multi-modal encoder. Second, we encourage the model to focus more on hard but diverse examples by proposing a focal loss for the image-text contrastive learning (ITC) objective, which alleviates the inherent limitations of overfitting and bias issues. Last, we perform multi-modal data augmentations for self-supervised learning via mining various examples by masking texts and rendering distortions on images. We show that the combination of these three innovations is effective for learning a pretrained model, leading to outstanding performance on multiple vision-language downstream tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval

    cs.CV 2025-10 conditional novelty 5.0 of 10

    MSAM introduces two drone-video/text datasets and a CLIP-based multi-semantic pooling model that reports 0.6–3.8 point R@1 gains over earlier video-text retrieval methods.

Pith tools