Pith. sign in

REVIEW 11 cited by

Understanding Dimensional Collapse in Contrastive Self-supervised Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.09348 v3 pith:O7BGHDWP submitted 2021-10-18 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords collapselearningcontrastiveembeddingdimensionalmethodsvectorsbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised visual representation learning aims to learn useful representations without relying on human annotations. Joint embedding approach bases on maximizing the agreement between embedding vectors from different views of the same image. Various methods have been proposed to solve the collapsing problem where all embedding vectors collapse to a trivial constant solution. Among these methods, contrastive learning prevents collapse via negative sample pairs. It has been shown that non-contrastive methods suffer from a lesser collapse problem of a different nature: dimensional collapse, whereby the embedding vectors end up spanning a lower-dimensional subspace instead of the entire available embedding space. Here, we show that dimensional collapse also happens in contrastive learning. In this paper, we shed light on the dynamics at play in contrastive learning that leads to dimensional collapse. Inspired by our theory, we propose a novel contrastive learning method, called DirectCLR, which directly optimizes the representation space without relying on an explicit trainable projector. Experiments show that DirectCLR outperforms SimCLR with a trainable linear projector on ImageNet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Between-User Collapse Under Popularity-Biased Feedback: A Centered-Covariance Theorem and Computable Phase Boundary

    cs.IR 2026-08 conditional novelty 7.0 of 10

    The centered user covariance under popularity-biased BPR training converges to a steady state proportional to the item-noise covariance, with a computable contraction/expansion boundary.

  2. SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation

    cs.IR 2026-07 conditional novelty 6.0 of 10

    SpecFormer is a spectral-aware Transformer that flattens the singular-value spectrum of embeddings to prevent embedding/attention collapse, outperforming baselines on CTR benchmarks and scaling with layer depth.

  3. Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Distilling TITAN and CARE slide embeddings into MIL aggregators gives reusable pretrained weights that beat from-scratch training on most of 15 pathology tasks, with the largest gains in few-shot and linear-probing settings.

  4. DeInfoReg: A Decoupled Learning Framework for Better Training Throughput

    cs.LG 2025-06 conditional novelty 6.0 of 10

    DeInfoReg trains deep networks with per-module local losses so gradients flow only within each module, improving accuracy and enabling pipeline parallelism, with speedups of up to 1.47x over single-GPU backpropagation.

  5. Visual Pre-Training on Unlabeled Images using Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Casting image-crop consistency as temporal-difference value learning improves visual representations on unlabeled web, scene, and video data.

  6. On the Importance of Embedding Norms in Self-Supervised Learning

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Embedding norms in self-supervised learning are not normalized away: they gate gradient size and encode confidence, so controlling them speeds up training.

  7. Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions

    cs.LG 2026-07 conditional novelty 5.0 of 10

    GSR uses class-adaptive spherical mixup to recondition the Gram matrix in recursive-least-squares continual learning, improving long-tailed accuracy by up to ~17 points while retaining O(D) cost.

  8. Denoising-While-Completing Network (DWCNet): Robust Point Cloud Completion Under Corruption

    cs.CV 2025-07 reject novelty 5.0 of 10

    A new corrupted point cloud completion benchmark and a network with contrastive feature filtering report top scores after fine-tuning, yet exhibit unexplained catastrophic failures on two corruptions.

  9. When Graph Contrastive Learning Backfires: Spectral Vulnerability and Defense in Recommendation

    cs.IR 2025-07 conditional novelty 4.0 of 10

    Graph contrastive learning smooths the embedding spectrum, which makes recommender systems more vulnerable to targeted item promotion attacks, and a spectral defense (SIM) can suppress those attacks.

  10. Generalized Category Discovery via Token Manifold Capacity Learning

    cs.LG 2025-05 conditional novelty 4.0 of 10

    MTMC adds a nuclear-norm loss on class tokens to GCD objectives and reports small accuracy gains on several image benchmarks.

  11. Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment

    cs.CV 2025-11 reject novelty 3.0 of 10

    JWTH achieves modest tissue-classification gains by adding attention pooling and stain augmentation to a DINOv3 backbone, but the biomarker claims in the abstract are unsupported by the experiments.

Pith tools