Pith. sign in

REVIEW 5 cited by

Dual-Path Convolutional Image-Text Embeddings with Instance Loss

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.05535 v4 pith:CQF2TUTO submitted 2017-11-15 cs.CV cs.MM

classification cs.CVcs.MM
keywords lossimagetextinstancelearnnetworkrankingapply
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Matching images and sentences demands a fine understanding of both modalities. In this paper, we propose a new system to discriminatively embed the image and text to a shared visual-textual space. In this field, most existing works apply the ranking loss to pull the positive image / text pairs close and push the negative pairs apart from each other. However, directly deploying the ranking loss is hard for network learning, since it starts from the two heterogeneous features to build inter-modal relationship. To address this problem, we propose the instance loss which explicitly considers the intra-modal data distribution. It is based on an unsupervised assumption that each image / text group can be viewed as a class. So the network can learn the fine granularity from every image/text group. The experiment shows that the instance loss offers better weight initialization for the ranking loss, so that more discriminative embeddings can be learned. Besides, existing works usually apply the off-the-shelf features, i.e., word2vec and fixed visual feature. So in a minor contribution, this paper constructs an end-to-end dual-path convolutional network to learn the image and text representations. End-to-end learning allows the system to directly learn from the data and fully utilize the supervision. On two generic retrieval datasets (Flickr30k and MSCOCO), experiments demonstrate that our method yields competitive accuracy compared to state-of-the-art methods. Moreover, in language based person retrieval, we improve the state of the art by a large margin. The code has been made publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large-scale Tag-based Font Retrieval with Generative Feature Learning

    cs.CV 2019-09 conditional novelty 6.0 of 10

    A generative feature learning and attention-based recognition-retrieval model improves tag-based font retrieval on a new 20,000-font benchmark dataset.

  2. Adversarial Representation Learning for Text-to-Image Matching

    cs.CV 2019-08 conditional novelty 6.0 of 10

    TIMAM combines an adversarial modality discriminator, norm-softmax identification losses, and a cross-modal projection matching loss with a BERT plus bidirectional LSTM text encoder, achieving state-of-the-art text-to...

  3. Matching Images and Text with Multi-modal Tensor Fusion and Re-ranking

    cs.CV 2019-08 conditional novelty 5.0 of 10

    A tensor-fusion network with cross-modal re-ranking achieves state-of-the-art image-text matching recall on Flickr30k and MSCOCO.

  4. Attract or Distract: Exploit the Margin of Open Set

    cs.CV 2019-08 conditional novelty 5.0 of 10

    Semantic Categorical Alignment and Semantic Contrastive Mapping improve open set domain adaptation by separating known and unknown classes in feature space.

  5. Do Cross Modal Systems Leverage Semantic Relationships?

    cs.CV 2019-09 reject novelty 4.0 of 10

    The authors introduce SemanticMap, a cosine-similarity based evaluation metric for cross-modal retrieval, and a single-stream network that encodes text as images, but the metric can be trivially gamed by collapsing em...

Pith tools