Pith. sign in

REVIEW 4 cited by

Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.00412 v2 pith:FDJMRD54 submitted 2021-09-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords multimodalfusioninformationinputresultstaskunimodalwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In multimodal sentiment analysis (MSA), the performance of a model highly depends on the quality of synthesized embeddings. These embeddings are generated from the upstream process called multimodal fusion, which aims to extract and combine the input unimodal raw data to produce a richer multimodal representation. Previous work either back-propagates the task loss or manipulates the geometric property of feature spaces to produce favorable fusion results, which neglects the preservation of critical task-related information that flows from input to the fusion results. In this work, we propose a framework named MultiModal InfoMax (MMIM), which hierarchically maximizes the Mutual Information (MI) in unimodal input pairs (inter-modality) and between multimodal fusion result and unimodal input in order to maintain task-related information through multimodal fusion. The framework is jointly trained with the main task (MSA) to improve the performance of the downstream MSA task. To address the intractable issue of MI bounds, we further formulate a set of computationally simple parametric and non-parametric methods to approximate their truth value. Experimental results on the two widely used datasets demonstrate the efficacy of our approach. The implementation of this work is publicly available at https://github.com/declare-lab/Multimodal-Infomax.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A dual-stream salience-context calibration that turns audio/visual signals into LLM-readable sentiment tokens improves sentiment classification on MOSI, MOSEI, CH-SIMS, and CH-SIMS v2.

  2. EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations

    cs.MM 2025-05 conditional novelty 6.0 of 10

    EmotionTalk provides 19,250 utterances from 744 Chinese dyadic dialogues with emotion, sentiment, and speaking-style caption annotations.

  3. Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis

    cs.MM 2026-05 unverdicted novelty 5.0 of 10

    MCAF reports state-of-the-art Acc-2/F1 on CMU-MOSI (86.52/86.51) and CMU-MOSEI (86.72/86.65), but its diffusion-denoising module is never specified in the methods.

  4. Decoding Visual Neural Representations by Multimodal with Dynamic Balancing

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A multimodal EEG-image-text contrastive framework with dynamic gradient balancing and stochastic noise improves zero-shot object recognition from EEG on ThingsEEG, raising top-1 accuracy from 13.8% to 15.8%.

Pith tools