Pith. sign in

REVIEW 3 cited by

Unicom: Universal and Compact Representation Learning for Image Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.05884 v1 pith:ISKBRD62 submitted 2023-04-12 cs.CV

classification cs.CV
keywords featurepre-trainedclassesimagepartialrepresentationretrievalcompact
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern image retrieval methods typically rely on fine-tuning pre-trained encoders to extract image-level descriptors. However, the most widely used models are pre-trained on ImageNet-1K with limited classes. The pre-trained feature representation is therefore not universal enough to generalize well to the diverse open-world classes. In this paper, we first cluster the large-scale LAION400M into one million pseudo classes based on the joint textual and visual features extracted by the CLIP model. Due to the confusion of label granularity, the automatically clustered dataset inevitably contains heavy inter-class conflict. To alleviate such conflict, we randomly select partial inter-class prototypes to construct the margin-based softmax loss. To further enhance the low-dimensional feature representation, we randomly select partial feature dimensions when calculating the similarities between embeddings and class-wise prototypes. The dual random partial selections are with respect to the class dimension and the feature dimension of the prototype matrix, making the classification conflict-robust and the feature embedding compact. Our method significantly outperforms state-of-the-art unsupervised and supervised image retrieval approaches on multiple benchmarks. The code and pre-trained models are released to facilitate future research https://github.com/deepglint/unicom.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Illuminating Visual Identity in Universal Multimodal Embeddings

    cs.CV 2026-08 conditional novelty 6.0 of 10

    By adding identity-aware sampling and a contrastive loss on a new 28-dataset benchmark, the authors build multimodal embeddings that are far better at visual identity matching without losing general retrieval accuracy.

  2. Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

    cs.CV 2025-09 conditional novelty 6.0 of 10

    GA-DMS with the WebPerson dataset sets new state-of-the-art Rank-1 accuracy on CUHK-PEDES, ICFG-PEDES, and RSTPReid.

  3. Locality-Sensitive Hashing for Efficient Hard Negative Sampling in Contrastive Learning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Binary LSH codes found by Hamming distance provide hard negatives for supervised contrastive learning at a fraction of the compute cost of exact pre-epoch sampling, with comparable or better accuracy on six benchmarks.

Pith tools