Pith. sign in

REVIEW 2 cited by

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.05804 v2 pith:NCATGAUS submitted 2023-10-09 cs.AI cs.CLcs.CVcs.MM

classification cs.AIcs.CLcs.CVcs.MM
keywords multimodalrepresentationadaptivehyper-modalityalmtanalysisaudioeffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales. With the obtained hyper-modality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA. In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis

    cs.LG 2024-12 conditional novelty 5.0 of 10

    DLF improves multimodal sentiment analysis by disentangling shared and specific features and steering cross-modal attention toward the dominant language modality.

  2. SentiXRL: An advanced large language Model Framework for Multilingual Fine-Grained Emotion Classification in Complex Text Environment

    cs.CL 2024-11 reject novelty 5.0 of 10

    SentiXRL is an LLM prompting and self-negotiation framework claimed to improve fine-grained emotion classification on Chinese and English benchmarks, but reported gains are small and internally inconsistent.

Pith tools