Pith. sign in

REVIEW 2 cited by

Multi-Modal Emotion recognition on IEMOCAP Dataset using Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.05788 v3 pith:LTF2DPKS submitted 2018-04-16 cs.AI cs.CVcs.HC

classification cs.AIcs.CVcs.HC
keywords emotiondataiemocaprecognitiondatasetresearchdetectionnetworks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Emotion recognition has become an important field of research in Human Computer Interactions as we improve upon the techniques for modelling the various aspects of behaviour. With the advancement of technology our understanding of emotions are advancing, there is a growing need for automatic emotion recognition systems. One of the directions the research is heading is the use of Neural Networks which are adept at estimating complex functions that depend on a large number and diverse source of input data. In this paper we attempt to exploit this effectiveness of Neural networks to enable us to perform multimodal Emotion recognition on IEMOCAP dataset using data from Speech, Text, and Motion capture data from face expressions, rotation and hand movements. Prior research has concentrated on Emotion detection from Speech on the IEMOCAP dataset, but our approach is the first that uses the multiple modes of data offered by IEMOCAP for a more robust and accurate emotion detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EmoTech: A Multi-modal Speech Emotion Recognition Using Multi-source Low-level Information with Hybrid Recurrent Network

    eess.AS 2025-01 conditional novelty 4.0 of 10

    A hybrid BiLSTM-CNN model using audio MFCCs and text embeddings reports 83.52% accuracy on five IEMOCAP emotion classes.

  2. A Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language

    cs.HC 2024-12 reject novelty 3.0 of 10

    A four-modality emotion recognition system that fuses facial, body, speech, and language cues is reported to reach 96.43 percent accuracy in a simulated, self-reported webcam test.

Pith tools