REVIEW 6 cited by
Scaling Language-Free Visual Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Visual Self-Supervised Learning (SSL) currently underperforms Contrastive Language-Image Pretraining (CLIP) in multimodal settings such as Visual Question Answering (VQA). This multimodal gap is often attributed to the semantics introduced by language supervision, even though visual SSL and CLIP models are often trained on different data. In this work, we ask the question: "Do visual self-supervised approaches lag behind CLIP due to the lack of language supervision, or differences in the training data?" We study this question by training both visual SSL and CLIP models on the same MetaCLIP data, and leveraging VQA as a diverse testbed for vision encoders. In this controlled setup, visual SSL models scale better than CLIP models in terms of data and model capacity, and visual SSL performance does not saturate even after scaling up to 7B parameters. Consequently, we observe visual SSL methods achieve CLIP-level performance on a wide range of VQA and classic vision benchmarks. These findings demonstrate that pure visual SSL can match language-supervised visual pretraining at scale, opening new opportunities for vision-centric representation learning.
Forward citations
Cited by 6 Pith papers
-
Meta CLIP 2: A Worldwide Scaling Recipe
A data curation and training recipe that scales CLIP from English-only data to 300+ languages from scratch, breaking the curse of multilinguality at ViT-H/14 scale.
-
Patch Policy: Efficient Embodied Control via Dense Visual Representations
Patch Policy shows that frozen dense ViT patch tokens, consumed through a block-causal attention mask, let lightweight robot policies beat pooled-feature policies and even a fine-tuned 7B vision-language-action model.
-
Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.
-
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
Dense scaling-law fits across model sizes, datasets and tasks show that MaMMUT (contrastive plus captioning loss) outperforms standard CLIP at large compute scales, with a consistent crossover around 1e10 to 1e11 GFLOPs.
-
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.
-
Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment
JWTH achieves modest tissue-classification gains by adding attention pooling and stain augmentation to a DINOv3 backbone, but the biomarker claims in the abstract are unsupported by the experiments.
Discussion (0). Continue with ORCID to comment.