REVIEW 3 cited by
ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different modalities as input and models the relationship between them. The model is pre-trained on four tasks simultaneously: Masked Language Modeling (MLM), Masked Object Classification (MOC), Masked Region Feature Regression (MRFR), and Image Text Matching (ITM). To further enhance the pre-training quality, we have collected a Large-scale weAk-supervised Image-Text (LAIT) dataset from Web. We first pre-train the model on this dataset, then conduct a second stage pre-training on Conceptual Captions and SBU Captions. Our experiments show that multi-stage pre-training strategy outperforms single-stage pre-training. We also fine-tune and evaluate our pre-trained ImageBERT model on image retrieval and text retrieval tasks, and achieve new state-of-the-art results on both MSCOCO and Flickr30k datasets.
Forward citations
Cited by 3 Pith papers
-
Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples
A vision-language pretraining method generates token-level negative image samples from a visual dictionary and combines them with textual negatives to improve fine-grained understanding.
-
GeoMM: On Geodesic Perspective for Multi-modal Learning
Graph shortest-path geodesic distance as a replacement for cosine similarity improves image-text contrastive pre-training by 1 to 3 retrieval points on ALBEF, TCL, and MAFA baselines.
-
CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance
Adding nearest-neighbor and cross nearest-neighbor supervision from frozen pretrained unimodal encoders to the CLIP loss improves lightweight vision-language models on zero-shot and retrieval benchmarks.
Discussion (0). Continue with ORCID to comment.