REVIEW 11 cited by
Understanding Dimensional Collapse in Contrastive Self-supervised Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Self-supervised visual representation learning aims to learn useful representations without relying on human annotations. Joint embedding approach bases on maximizing the agreement between embedding vectors from different views of the same image. Various methods have been proposed to solve the collapsing problem where all embedding vectors collapse to a trivial constant solution. Among these methods, contrastive learning prevents collapse via negative sample pairs. It has been shown that non-contrastive methods suffer from a lesser collapse problem of a different nature: dimensional collapse, whereby the embedding vectors end up spanning a lower-dimensional subspace instead of the entire available embedding space. Here, we show that dimensional collapse also happens in contrastive learning. In this paper, we shed light on the dynamics at play in contrastive learning that leads to dimensional collapse. Inspired by our theory, we propose a novel contrastive learning method, called DirectCLR, which directly optimizes the representation space without relying on an explicit trainable projector. Experiments show that DirectCLR outperforms SimCLR with a trainable linear projector on ImageNet.
Forward citations
Cited by 11 Pith papers
-
Between-User Collapse Under Popularity-Biased Feedback: A Centered-Covariance Theorem and Computable Phase Boundary
The centered user covariance under popularity-biased BPR training converges to a steady state proportional to the item-noise covariance, with a computable contraction/expansion boundary.
-
SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation
SpecFormer is a spectral-aware Transformer that flattens the singular-value spectrum of embeddings to prevent embedding/attention collapse, outperforming baselines on CTR benchmarks and scaling with layer depth.
-
Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models
Distilling TITAN and CARE slide embeddings into MIL aggregators gives reusable pretrained weights that beat from-scratch training on most of 15 pathology tasks, with the largest gains in few-shot and linear-probing settings.
-
DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
DeInfoReg trains deep networks with per-module local losses so gradients flow only within each module, improving accuracy and enabling pipeline parallelism, with speedups of up to 1.47x over single-GPU backpropagation.
-
Visual Pre-Training on Unlabeled Images using Reinforcement Learning
Casting image-crop consistency as temporal-difference value learning improves visual representations on unlabeled web, scene, and video data.
-
On the Importance of Embedding Norms in Self-Supervised Learning
Embedding norms in self-supervised learning are not normalized away: they gate gradient size and encode confidence, so controlling them speeds up training.
-
Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions
GSR uses class-adaptive spherical mixup to recondition the Gram matrix in recursive-least-squares continual learning, improving long-tailed accuracy by up to ~17 points while retaining O(D) cost.
-
Denoising-While-Completing Network (DWCNet): Robust Point Cloud Completion Under Corruption
A new corrupted point cloud completion benchmark and a network with contrastive feature filtering report top scores after fine-tuning, yet exhibit unexplained catastrophic failures on two corruptions.
-
When Graph Contrastive Learning Backfires: Spectral Vulnerability and Defense in Recommendation
Graph contrastive learning smooths the embedding spectrum, which makes recommender systems more vulnerable to targeted item promotion attacks, and a spectral defense (SIM) can suppress those attacks.
-
Generalized Category Discovery via Token Manifold Capacity Learning
MTMC adds a nuclear-norm loss on class tokens to GCD objectives and reports small accuracy gains on several image benchmarks.
-
Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment
JWTH achieves modest tissue-classification gains by adding attention pooling and stain augmentation to a DINOv3 backbone, but the biomarker claims in the abstract are unsupported by the experiments.
Discussion (0). Continue with ORCID to comment.