REVIEW 4 cited by
What to align in multimodal contrastive learning?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Humans perceive the world through multisensory integration, blending the information of different modalities to adapt their behavior. Contrastive learning offers an appealing solution for multimodal self-supervised learning. Indeed, by considering each modality as a different view of the same entity, it learns to align features of different modalities in a shared representation space. However, this approach is intrinsically limited as it only learns shared or redundant information between modalities, while multimodal interactions can arise in other ways. In this work, we introduce CoMM, a Contrastive MultiModal learning strategy that enables the communication between modalities in a single multimodal space. Instead of imposing cross- or intra- modality constraints, we propose to align multimodal representations by maximizing the mutual information between augmented versions of these multimodal features. Our theoretical analysis shows that shared, synergistic and unique terms of information naturally emerge from this formulation, allowing us to estimate multimodal interactions beyond redundancy. We test CoMM both in a controlled and in a series of real-world settings: in the former, we demonstrate that CoMM effectively captures redundant, unique and synergistic information between modalities. In the latter, CoMM learns complex multimodal interactions and achieves state-of-the-art results on the seven multimodal benchmarks. Code is available at https://github.com/Duplums/CoMM
Forward citations
Cited by 4 Pith papers
-
Hyperbolic Multimodal Continual Learning
The paper proposes projecting continual updates away from old-task spatial directions in hyperbolic multimodal models, claims this is theoretically required to prevent forgetting, but the necessity claim is unproven a...
-
Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion
A regularization method enforces diverse intra-modal embeddings and bounded inter-modal drift to improve both multimodal fusion and unimodal robustness.
-
Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach
A dual-path ResNet/DenseNet framework with multi-stage contrastive learning and confidence-driven gradient modulation is presented for multimodal human activity recognition, with reported improvements on four public datasets.
-
I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts
I2MoE improves multimodal fusion by training interaction-specialized experts with perturbed-modality supervision and reweighting their outputs per sample.
Discussion (0). Continue with ORCID to comment.