A CNN-LSTM model with MFCC features classifies seven speech emotions on a SAVEE/RAVDESS subset at 61.07% accuracy, with anger (75.31%) and neutral (71.70%) recognized best.
Regarding the feature extraction, features from four aspects are adopted for further analysis and the MFCC processing provided by the Librosa is specifically vital
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture
A CNN-LSTM model with MFCC features classifies seven speech emotions on a SAVEE/RAVDESS subset at 61.07% accuracy, with anger (75.31%) and neutral (71.70%) recognized best.