Pith. sign in

REVIEW 1 cited by

Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19560 v1 pith:SHIBXQTQ submitted 2024-10-25 cs.CV cs.AIcs.LGeess.IVeess.SP

classification cs.CVcs.AIcs.LGeess.IVeess.SP
keywords learningarchitecturec-jepajoint-embeddingpredictivevisualcollapseentire
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent advancements in unsupervised visual representation learning, the Joint-Embedding Predictive Architecture (JEPA) has emerged as a significant method for extracting visual features from unlabeled imagery through an innovative masking strategy. Despite its success, two primary limitations have been identified: the inefficacy of Exponential Moving Average (EMA) from I-JEPA in preventing entire collapse and the inadequacy of I-JEPA prediction in accurately learning the mean of patch representations. Addressing these challenges, this study introduces a novel framework, namely C-JEPA (Contrastive-JEPA), which integrates the Image-based Joint-Embedding Predictive Architecture with the Variance-Invariance-Covariance Regularization (VICReg) strategy. This integration is designed to effectively learn the variance/covariance for preventing entire collapse and ensuring invariance in the mean of augmented views, thereby overcoming the identified limitations. Through empirical and theoretical evaluations, our work demonstrates that C-JEPA significantly enhances the stability and quality of visual representation learning. When pre-trained on the ImageNet-1K dataset, C-JEPA exhibits rapid and improved convergence in both linear probing and fine-tuning performance metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Simplifying DINO via Coding Rate Regularization

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Replacing DINO's complex anti-collapse machinery with an explicit coding rate regularizer yields simpler, more stable, and higher-performing self-supervised models.

Pith tools