Pith. sign in

REVIEW 3 cited by

Improving fine-grained understanding in image-text pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.09865 v1 pith:UWAX6BIV submitted 2024-01-18 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords fine-grainedimageinformationpatcheslosssparctokencontrastive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce SPARse Fine-grained Contrastive Alignment (SPARC), a simple method for pretraining more fine-grained multimodal representations from image-text pairs. Given that multiple image patches often correspond to single words, we propose to learn a grouping of image patches for every token in the caption. To achieve this, we use a sparse similarity metric between image patches and language tokens and compute for each token a language-grouped vision embedding as the weighted average of patches. The token and language-grouped vision embeddings are then contrasted through a fine-grained sequence-wise loss that only depends on individual samples and does not require other batch samples as negatives. This enables more detailed information to be learned in a computationally inexpensive manner. SPARC combines this fine-grained loss with a contrastive loss between global image and text embeddings to learn representations that simultaneously encode global and local information. We thoroughly evaluate our proposed method and show improved performance over competing approaches both on image-level tasks relying on coarse-grained information, e.g. classification, as well as region-level tasks relying on fine-grained information, e.g. retrieval, object detection, and segmentation. Moreover, SPARC improves model faithfulness and captioning in foundational vision-language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. QuARI: Query Adaptive Retrieval Improvement

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A hypernetwork predicts a query-specific low-rank linear projection that reshapes frozen VLM embeddings, improving retrieval on ILIAS and INQUIRE.

  2. AGA: An adaptive group alignment framework for structured medical cross-modal representation learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    AGA builds token-to-patch and patch-to-token groups via a sparse similarity matrix with adaptive thresholds, and trains a medical image-text model with within-pair contrastive losses.

  3. Xray-Visual Models: Scaling Vision models on Industry Scale Data

    cs.CV 2026-02 conditional novelty 4.0 of 10

    A 2-billion-parameter vision encoder trained on 15B+ image-text and billions of video-hashtag pairs reports SOTA ImageNet linear-probe, Kinetics, and retrieval numbers, but relies on proprietary data and has several v...

Pith tools