Pith. sign in

REVIEW 6 cited by

Scaling Language-Free Visual Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.01017 v1 pith:TOKO2W37 submitted 2025-04-01 cs.CV

classification cs.CV
keywords visualclipdatamodelslearningquestionevenlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual Self-Supervised Learning (SSL) currently underperforms Contrastive Language-Image Pretraining (CLIP) in multimodal settings such as Visual Question Answering (VQA). This multimodal gap is often attributed to the semantics introduced by language supervision, even though visual SSL and CLIP models are often trained on different data. In this work, we ask the question: "Do visual self-supervised approaches lag behind CLIP due to the lack of language supervision, or differences in the training data?" We study this question by training both visual SSL and CLIP models on the same MetaCLIP data, and leveraging VQA as a diverse testbed for vision encoders. In this controlled setup, visual SSL models scale better than CLIP models in terms of data and model capacity, and visual SSL performance does not saturate even after scaling up to 7B parameters. Consequently, we observe visual SSL methods achieve CLIP-level performance on a wide range of VQA and classic vision benchmarks. These findings demonstrate that pure visual SSL can match language-supervised visual pretraining at scale, opening new opportunities for vision-centric representation learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Meta CLIP 2: A Worldwide Scaling Recipe

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A data curation and training recipe that scales CLIP from English-only data to 300+ languages from scratch, breaking the curse of multilinguality at ViT-H/14 scale.

  2. Patch Policy: Efficient Embodied Control via Dense Visual Representations

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Patch Policy shows that frozen dense ViT patch tokens, consumed through a block-causal attention mask, let lightweight robot policies beat pooled-feature policies and even a fine-tuned 7B vision-language-action model.

  3. Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning

    cs.CV 2025-06 reject novelty 6.0 of 10

    AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.

  4. Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Dense scaling-law fits across model sizes, datasets and tasks show that MaMMUT (contrastive plus captioning loss) outperforms standard CLIP at large compute scales, with a consistent crossover around 1e10 to 1e11 GFLOPs.

  5. SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.

  6. Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment

    cs.CV 2025-11 reject novelty 3.0 of 10

    JWTH achieves modest tissue-classification gains by adding attention pooling and stain augmentation to a DINOv3 backbone, but the biomarker claims in the abstract are unsupported by the experiments.

Pith tools