REVIEW 5 cited by
Multi-modal Attention for Speech Emotion Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Emotion represents an essential aspect of human speech that is manifested in speech prosody. Speech, visual, and textual cues are complementary in human communication. In this paper, we study a hybrid fusion method, referred to as multi-modal attention network (MMAN) to make use of visual and textual cues in speech emotion recognition. We propose a novel multi-modal attention mechanism, cLSTM-MMA, which facilitates the attention across three modalities and selectively fuse the information. cLSTM-MMA is fused with other uni-modal sub-networks in the late fusion. The experiments show that speech emotion recognition benefits significantly from visual and textual cues, and the proposed cLSTM-MMA alone is as competitive as other fusion methods in terms of accuracy, but with a much more compact network structure. The proposed hybrid network MMAN achieves state-of-the-art performance on IEMOCAP database for emotion recognition.
Forward citations
Cited by 5 Pith papers
-
Image Editing As Programs with Diffusion Models
IEAP decomposes complex editing instructions into atomic operations executed sequentially on a diffusion transformer, and reports state-of-the-art results on MagicBrush and AnyEdit.
-
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
OmniConsistency is a style-agnostic consistency module for Flux that preserves structure and details during stylization with arbitrary LoRAs, reaching GPT-4o-level content consistency.
-
WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending
A text-to-image pipeline that lets users restyle individual characters or regions of artistic typography and iteratively refine them with region-specific prompts.
-
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
A decoupled-attention adapter transfers image-pair edits to new photos in diffusion transformers, trained with a new 218-task visual editing dataset.
-
Learning Annotation Consensus for Continuous Emotion Recognition
A consensus network over multiple annotator labels improves continuous emotion prediction on RECOLA, but the claimed COGNIMUSE results are absent from the paper.
Discussion (0). Continue with ORCID to comment.