Pith. sign in

REVIEW 3 cited by

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.11256 v1 pith:7KJEZDGQ submitted 2022-11-21 cs.CL

classification cs.CL
keywords sentimentemotionmultimodalanalysisemotionsperiodrecognitionsentiments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a longer period. However, most existing works study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two. In this paper, we propose a multimodal sentiment knowledge-sharing framework (UniMSE) that unifies MSA and ERC tasks from features, labels, and models. We perform modality fusion at the syntactic and semantic levels and introduce contrastive learning between modalities and samples to better capture the difference and consistency between sentiments and emotions. Experiments on four public benchmark datasets, MOSI, MOSEI, MELD, and IEMOCAP, demonstrate the effectiveness of the proposed method and achieve consistent improvements compared with state-of-the-art methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A dual-stream salience-context calibration that turns audio/visual signals into LLM-readable sentiment tokens improves sentiment classification on MOSI, MOSEI, CH-SIMS, and CH-SIMS v2.

  2. Partitioner Guided Modal Learning Framework

    cs.CL 2025-07 conditional novelty 6.0 of 10

    PgM segments multimodal representations into uni-modal and paired-modal features with cumulative-softmax gates and trains them with separate learners, reconstruction, and classification losses, yielding accuracy gains...

  3. Causal Emotion Recognition in Conversation: Context Saturation and Discourse-Marker Evidence

    cs.CL 2026-01 conditional novelty 5.0 of 10

    Using only past turns, ERC accuracy saturates within 10–30 preceding utterances; hierarchical encoding and SenticNet add little once context is present, and Sad turns show the largest context benefit and fewer left-pe...

Pith tools