REVIEW 6 cited by
Self-supervised Learning from a Multi-view Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As a subset of unsupervised representation learning, self-supervised representation learning adopts self-defined signals as supervision and uses the learned representation for downstream tasks, such as object detection and image captioning. Many proposed approaches for self-supervised learning follow naturally a multi-view perspective, where the input (e.g., original images) and the self-supervised signals (e.g., augmented images) can be seen as two redundant views of the data. Building from this multi-view perspective, this paper provides an information-theoretical framework to better understand the properties that encourage successful self-supervised learning. Specifically, we demonstrate that self-supervised learned representations can extract task-relevant information and discard task-irrelevant information. Our theoretical framework paves the way to a larger space of self-supervised learning objective design. In particular, we propose a composite objective that bridges the gap between prior contrastive and predictive learning objectives, and introduce an additional objective term to discard task-irrelevant information. To verify our analysis, we conduct controlled experiments to evaluate the impact of the composite objectives. We also explore our framework's empirical generalization beyond the multi-view perspective, where the cross-view redundancy may not be clearly observed.
Forward citations
Cited by 6 Pith papers
-
Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised Learning
S2DC combines dual-softmax patch correspondence with Sharpe-ratio-weighted structural consistency to learn structure-aware representations for 3D medical image self-supervised learning.
-
Automatically Identify and Rectify: Robust Deep Contrastive Multi-view Clustering in Noisy Scenarios
Multi-view clustering with GMM-based noise identification, first-view-anchored hybrid rectification, and soft-label-gated contrastive learning beats 11 baselines on six noisy benchmarks.
-
A Cross Modal Knowledge Distillation & Data Augmentation Recipe for Improving Transcriptomics Representations through Morphological Features
A CLIP-style distillation from microscopy to transcriptomics, combined with a perturbation-embedding augmentation, improves biological relationship recall on out-of-distribution datasets while preserving interpretability.
-
Aligning Multimodal Representations through an Information Bottleneck
A regularizer derived from an information-bottleneck bound, essentially a mean-squared alignment loss, reduces modality-specific information and improves multimodal alignment and image captioning.
-
On the Transferability and Discriminability of Repersentation Learning in Unsupervised Domain Adaptation
RLGLC combines a relaxed Wasserstein alignment with a contrastive local consistency term for UDA and reports SOTA results, but the proof that such a term is necessary is not rigorous.
-
Task-Oriented Low-Label Semantic Communication With Self-Supervised Learning
SLSCom pre-trains a semantic encoder with contrastive and reconstruction pretext tasks on unlabeled data, then jointly fine-tunes it with JSCC and a classifier, improving accuracy under few labels and low SNR.
Discussion (0). Continue with ORCID to comment.