Pith. sign in

REVIEW 2 cited by

SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15308 v4 pith:QGK4FA3E submitted 2023-10-23 cs.CV cs.LG

classification cs.CVcs.LG
keywords clipsam-clipmodelmodelssemanticunderstandingvfmsvision
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The landscape of publicly available vision foundation models (VFMs), such as CLIP and Segment Anything Model (SAM), is expanding rapidly. VFMs are endowed with distinct capabilities stemming from their pre-training objectives. For instance, CLIP excels in semantic understanding, while SAM specializes in spatial understanding for segmentation. In this work, we introduce a simple recipe to efficiently merge VFMs into a unified model that absorbs their expertise. Our method integrates techniques of multi-task learning, continual learning, and distillation. Further, it demands significantly less computational cost compared to traditional multi-task training from scratch, and it only needs a small fraction of the pre-training datasets that were initially used to train individual models. By applying our method to SAM and CLIP, we obtain SAM-CLIP: a unified model that combines the capabilities of SAM and CLIP into a single vision transformer. Compared with deploying SAM and CLIP independently, our merged model, SAM-CLIP, reduces storage and compute costs for inference, making it well-suited for edge device applications. We show that SAM-CLIP not only retains the foundational strengths of SAM and CLIP, but also introduces synergistic functionalities, notably in zero-shot semantic segmentation, where SAM-CLIP establishes new state-of-the-art results on 5 benchmarks. It outperforms previous models that are specifically designed for this task by a large margin, including +6.8% and +5.9% mean IoU improvement on Pascal-VOC and COCO-Stuff datasets, respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A knowledge transfer framework that aggregates features and logits from over 100 heterogeneous pre-trained teacher models via sign voting and pseudo-class voting, improving unsupervised object recognition accuracy.

  2. Wearable Accelerometer Foundation Models for Health via Knowledge Distillation

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A knowledge-distilled accelerometry encoder, taught by an unsupervised PPG teacher on 20 million minutes of paired wearable data, predicts heart rate, heart-rate variability, demographics, and 46 health conditions fro...

Pith tools