Pith. sign in

REVIEW 6 cited by

Self-supervised Learning from a Multi-view Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.05576 v4 pith:VQHST3GE submitted 2020-06-10 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningself-supervisedmulti-viewperspectiveframeworkinformationobjectiverepresentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As a subset of unsupervised representation learning, self-supervised representation learning adopts self-defined signals as supervision and uses the learned representation for downstream tasks, such as object detection and image captioning. Many proposed approaches for self-supervised learning follow naturally a multi-view perspective, where the input (e.g., original images) and the self-supervised signals (e.g., augmented images) can be seen as two redundant views of the data. Building from this multi-view perspective, this paper provides an information-theoretical framework to better understand the properties that encourage successful self-supervised learning. Specifically, we demonstrate that self-supervised learned representations can extract task-relevant information and discard task-irrelevant information. Our theoretical framework paves the way to a larger space of self-supervised learning objective design. In particular, we propose a composite objective that bridges the gap between prior contrastive and predictive learning objectives, and introduce an additional objective term to discard task-irrelevant information. To verify our analysis, we conduct controlled experiments to evaluate the impact of the composite objectives. We also explore our framework's empirical generalization beyond the multi-view perspective, where the cross-view redundancy may not be clearly observed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 12 citations worldwide. Full citation record

  1. Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    S2DC combines dual-softmax patch correspondence with Sharpe-ratio-weighted structural consistency to learn structure-aware representations for 3D medical image self-supervised learning.

  2. Automatically Identify and Rectify: Robust Deep Contrastive Multi-view Clustering in Noisy Scenarios

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Multi-view clustering with GMM-based noise identification, first-view-anchored hybrid rectification, and soft-label-gated contrastive learning beats 11 baselines on six noisy benchmarks.

  3. A Cross Modal Knowledge Distillation & Data Augmentation Recipe for Improving Transcriptomics Representations through Morphological Features

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A CLIP-style distillation from microscopy to transcriptomics, combined with a perturbation-embedding augmentation, improves biological relationship recall on out-of-distribution datasets while preserving interpretability.

  4. Aligning Multimodal Representations through an Information Bottleneck

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A regularizer derived from an information-bottleneck bound, essentially a mean-squared alignment loss, reduces modality-specific information and improves multimodal alignment and image captioning.

  5. On the Transferability and Discriminability of Repersentation Learning in Unsupervised Domain Adaptation

    cs.CV 2025-05 reject novelty 5.0 of 10

    RLGLC combines a relaxed Wasserstein alignment with a contrastive local consistency term for UDA and reports SOTA results, but the proof that such a term is necessary is not rigorous.

  6. Task-Oriented Low-Label Semantic Communication With Self-Supervised Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    SLSCom pre-trains a semantic encoder with contrastive and reconstruction pretext tasks on unlabeled data, then jointly fine-tunes it with JSCC and a classifier, improving accuracy under few labels and low SNR.

Pith tools