REVIEW 7 cited by
On Mutual Information Maximization for Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Many recent methods for unsupervised or self-supervised representation learning train feature extractors by maximizing an estimate of the mutual information (MI) between different views of the data. This comes with several immediate problems: For example, MI is notoriously hard to estimate, and using it as an objective for representation learning may lead to highly entangled representations due to its invariance under arbitrary invertible transformations. Nevertheless, these methods have been repeatedly shown to excel in practice. In this paper we argue, and provide empirical evidence, that the success of these methods cannot be attributed to the properties of MI alone, and that they strongly depend on the inductive bias in both the choice of feature extractor architectures and the parametrization of the employed MI estimators. Finally, we establish a connection to deep metric learning and argue that this interpretation may be a plausible explanation for the success of the recently introduced methods.
Forward citations
Cited by 7 Pith papers
-
Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.
-
DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
DeInfoReg trains deep networks with per-module local losses so gradients flow only within each module, improving accuracy and enabling pipeline parallelism, with speedups of up to 1.47x over single-GPU backpropagation.
-
A Mathematical Perspective On Contrastive Learning
A probabilistic tilting framework for contrastive learning yields closed-form Gaussian results showing which conditional statistics each loss can recover.
-
Language-Aware Information Maximization for Transductive Few-Shot CLIP
LIMO, a transductive loss combining mutual information, zero-shot KL regularization, and LoRA, sets new state-of-the-art few-shot accuracy for CLIP on 11 datasets.
-
Structure Maintained Representation Learning Neural Network for Causal Inference
SMRLNN combines an adversarial discriminator with a canonical-correlation structure keeper to improve individual treatment effect estimation.
-
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
Contrastive Successor Features recover ground-truth RL states up to a linear map whenever the skill-conditioned transition differences follow a von Mises-Fisher distribution and policies are diverse.
-
C-LEAD: Contrastive Learning for Enhanced Adversarial Defense
Contrastive learning with adversarial perturbations as positive pairs improves robustness of ResNet models on CIFAR-10, but evidence is weakened by missing baselines and inconsistent reporting.
Discussion (0). Continue with ORCID to comment.