Pith. sign in

REVIEW 1 cited by

Improving Multimodal fusion via Mutual Dependency Maximisation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.00922 v2 pith:RQOCHXUS submitted 2021-08-31 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords multimodalfusionrepresentationsanalysisdatasetsdependencymodalitiespenalties
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Multimodal sentiment analysis is a trending area of research, and the multimodal fusion is one of its most active topic. Acknowledging humans communicate through a variety of channels (i.e visual, acoustic, linguistic), multimodal systems aim at integrating different unimodal representations into a synthetic one. So far, a consequent effort has been made on developing complex architectures allowing the fusion of these modalities. However, such systems are mainly trained by minimising simple losses such as $L_1$ or cross-entropy. In this work, we investigate unexplored penalties and propose a set of new objectives that measure the dependency between modalities. We demonstrate that our new penalties lead to a consistent improvement (up to $4.3$ on accuracy) across a large variety of state-of-the-art models on two well-known sentiment analysis datasets: \texttt{CMU-MOSI} and \texttt{CMU-MOSEI}. Our method not only achieves a new SOTA on both datasets but also produces representations that are more robust to modality drops. Finally, a by-product of our methods includes a statistical network which can be used to interpret the high dimensional representations learnt by the model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Asymmetric Reinforcing against Multi-modal Representation Bias

    cs.CV 2025-01 reject novelty 6.0 of 10

    ARM, a mutual-information-based asymmetric reinforcement method, narrows modality contribution gaps and reports improved accuracy on three multimodal classification datasets.

Pith tools