REVIEW 2 cited by
Improving Multimodal Accuracy Through Modality Pre-training and Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Training a multimodal network is challenging and it requires complex architectures to achieve reasonable performance. We show that one reason for this phenomena is the difference between the convergence rate of various modalities. We address this by pre-training modality-specific sub-networks in multimodal architectures independently before end-to-end training of the entire network. Furthermore, we show that the addition of an attention mechanism between sub-networks after pre-training helps identify the most important modality during ambiguous scenarios boosting the performance. We demonstrate that by performing these two tricks a simple network can achieve similar performance to a complicated architecture that is significantly more expensive to train on multiple tasks including sentiment analysis, emotion recognition, and speaker trait recognition.
Forward citations
Cited by 2 Pith papers
-
FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data
FiGuRO estimates the intrinsic dimensionality of shared and private subspaces in multi-modal data by adaptively growing or shrinking low-rank bottleneck layers guided by a reconstruction-fidelity budget.
-
Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning
DAAN combines a differential-attention module (QDMA) and a sample-level gradient modulation block (CSGM) to improve audio-visual zero-shot classification, reporting the best UCF101 GZSL harmonic mean so far.
Discussion (0). Continue with ORCID to comment.