Pith. sign in

REVIEW 3 cited by

Image-text Retrieval: A Survey on Recent Research and Development

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.14713 v3 pith:5CYV2VEC submitted 2022-03-28 cs.IR cs.AI

classification cs.IRcs.AI
keywords approachesresearchretrievalcross-modalfeatureimage-textmodalityperspective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the past few years, cross-modal image-text retrieval (ITR) has experienced increased interest in the research community due to its excellent research value and broad real-world application. It is designed for the scenarios where the queries are from one modality and the retrieval galleries from another modality. This paper presents a comprehensive and up-to-date survey on the ITR approaches from four perspectives. By dissecting an ITR system into two processes: feature extraction and feature alignment, we summarize the recent advance of the ITR approaches from these two perspectives. On top of this, the efficiency-focused study on the ITR system is introduced as the third perspective. To keep pace with the times, we also provide a pioneering overview of the cross-modal pre-training ITR approaches as the fourth perspective. Finally, we outline the common benchmark datasets and valuation metric for ITR, and conduct the accuracy comparison among the representative ITR approaches. Some critical yet less studied issues are discussed at the end of the paper.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A variance-based, retraining-free pruning framework for vision-language models that allocates per-layer sparsity and outperforms Wanda and SparseGPT at high sparsity.

  2. Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting

    cs.CV 2025-10 conditional novelty 5.0 of 10

    CAW adds a confidence-weighted KL loss and feature-alignment regularization to CLIP adversarial fine-tuning, raising average AutoAttack robust accuracy from 31.6% to 33.5% on 15 datasets.

  3. Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes

    cs.LG 2025-01 conditional novelty 5.0 of 10

    CM50, a 33k-meme dataset with GPT-4o-generated annotations, and mtrCLIP, a fine-tuned CLIP model, together improve meme-text retrieval on MemeCap over the original CLIP.

Pith tools