Pith. sign in

REVIEW 2 cited by

Improving Multimodal Accuracy Through Modality Pre-training and Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.06102 v1 pith:E4ZAITTQ submitted 2020-11-11 cs.AI

classification cs.AI
keywords multimodalnetworkperformancepre-trainingachievearchitecturesattentionmodality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training a multimodal network is challenging and it requires complex architectures to achieve reasonable performance. We show that one reason for this phenomena is the difference between the convergence rate of various modalities. We address this by pre-training modality-specific sub-networks in multimodal architectures independently before end-to-end training of the entire network. Furthermore, we show that the addition of an attention mechanism between sub-networks after pre-training helps identify the most important modality during ambiguous scenarios boosting the performance. We demonstrate that by performing these two tricks a simple network can achieve similar performance to a complicated architecture that is significantly more expensive to train on multiple tasks including sentiment analysis, emotion recognition, and speaker trait recognition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data

    cs.LG 2026-08 conditional novelty 6.0 of 10

    FiGuRO estimates the intrinsic dimensionality of shared and private subspaces in multi-modal data by adaptively growing or shrinking low-rank bottleneck layers guided by a reconstruction-fidelity budget.

  2. Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning

    cs.CV 2024-12 conditional novelty 4.0 of 10

    DAAN combines a differential-attention module (QDMA) and a sample-level gradient modulation block (CSGM) to improve audio-visual zero-shot classification, reporting the best UCF101 GZSL harmonic mean so far.

Pith tools