Pith. sign in

REVIEW 3 cited by

ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.07966 v2 pith:ZGPSCEC3 submitted 2020-01-22 cs.CV

classification cs.CV
keywords modelpre-trainingimage-textimagebertmaskedpre-trainedcaptionsdataset
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different modalities as input and models the relationship between them. The model is pre-trained on four tasks simultaneously: Masked Language Modeling (MLM), Masked Object Classification (MOC), Masked Region Feature Regression (MRFR), and Image Text Matching (ITM). To further enhance the pre-training quality, we have collected a Large-scale weAk-supervised Image-Text (LAIT) dataset from Web. We first pre-train the model on this dataset, then conduct a second stage pre-training on Conceptual Captions and SBU Captions. Our experiments show that multi-stage pre-training strategy outperforms single-stage pre-training. We also fine-tune and evaluate our pre-trained ImageBERT model on image retrieval and text retrieval tasks, and achieve new state-of-the-art results on both MSCOCO and Flickr30k datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A vision-language pretraining method generates token-level negative image samples from a visual dictionary and combines them with textual negatives to improve fine-grained understanding.

  2. GeoMM: On Geodesic Perspective for Multi-modal Learning

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Graph shortest-path geodesic distance as a replacement for cosine similarity improves image-text contrastive pre-training by 1 to 3 retrieval points on ALBEF, TCL, and MAFA baselines.

  3. CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Adding nearest-neighbor and cross nearest-neighbor supervision from frozen pretrained unimodal encoders to the CLIP loss improves lightweight vision-language models on zero-shot and retrieval benchmarks.

Pith tools