Pith. sign in

REVIEW 1 cited by

Gated Mechanism for Attention Based Multimodal Sentiment Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.01043 v1 pith:I6DVXFXN submitted 2020-02-21 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords multimodalsentimentanalysiscrosslearningmodalintensityinteractions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal sentiment analysis has recently gained popularity because of its relevance to social media posts, customer service calls and video blogs. In this paper, we address three aspects of multimodal sentiment analysis; 1. Cross modal interaction learning, i.e. how multiple modalities contribute to the sentiment, 2. Learning long-term dependencies in multimodal interactions and 3. Fusion of unimodal and cross modal cues. Out of these three, we find that learning cross modal interactions is beneficial for this problem. We perform experiments on two benchmark datasets, CMU Multimodal Opinion level Sentiment Intensity (CMU-MOSI) and CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) corpus. Our approach on both these tasks yields accuracies of 83.9% and 81.1% respectively, which is 1.6% and 1.34% absolute improvement over current state-of-the-art.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets

    cs.CV 2025-01 conditional novelty 6.0 of 10

    CM3T shows that multi-head vision adapters plus cross-attention adapters can adapt frozen supervised-pretrained video transformers with a fraction of the trainable parameters of full fine-tuning.

Pith tools