REVIEW 3 major objections 6 minor 98 references
Spatio-Temporal Fuzzy-oriented Multi-Modal Meta-Learning for Fine-grained Emotion Recognition
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A meta-learning framework with fuzzy rules and spatio-temporal encoding claims top accuracy for fine-grained emotion recognition, even with limited or noisy data.
desk verdict A genuinely integrated method whose two self-labeled benchmarks make the headline five-dataset claim weaker than it looks; the three external benchmarks still show real gains. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key objects are the two fuzzy inference systems: the fuzzy component inference system (FCIS), which converts facial component states into linguistic condition sequences using membership functions, and the fuzzy knowledge inference system (FKIS), which maps those sequences to emotional intensity levels (low, medium, high) for each of six basic emotions. These fuzzy systems serve both as a feature augmentation module inside the model and as a rule-based label generator for two new benchmarks, DISFA-FER and WFLW-FER. The learning machinery is a meta-recurrent neural network with a spatio-temporal convolutional encoder that performs bi-level optimization over tasks constructed from multi-modal views.
What would settle it
Take a new facial expression dataset with human-annotated fine-grained emotion and intensity labels (not derived from the paper's fuzzy rules), train ST-F2M with only the feature-augmentation role of the fuzzy systems, and compare against a version with the fuzzy modules removed. If the fuzzy-augmented model does not outperform the ablated version on this independently labeled dataset, the claim that fuzzy semantics improve emotion recognition would be falsified.
Extended reading notes
Core claim
The central claim is that fine-grained emotion recognition can be made data-efficient and robust by decomposing emotions into fuzzy components, encoding spatio-temporal heterogeneity with convolutions, and learning general emotion meta-knowledge via bi-level optimization. The paper introduces two fuzzy inference systems (FCIS and FKIS) that quantify emotional intensity and assign fuzzy semantic information to representations, and it uses a meta-recurrent neural network to achieve rapid adaptation. On five datasets, ST-F2M reports accuracy improvements of up to several points over previous methods, and it maintains high performance under synthetic noise and with only 20% of training data plus 30% mislabeled samples.
Load-bearing premise
The hand-coded fuzzy rules in FCIS and FKIS are assumed to be a valid compositional model of emotion, and the ground-truth labels for two of the five benchmark datasets are generated from those same rules, so the strong results on those datasets could reflect self-consistency rather than independent recognition.
Editorial extensions
If this is right
- If the claimed results hold, ST-F2M would enable fine-grained emotion recognition with far fewer annotated samples, since the fuzzy rules supply emotional intensity information without manual labeling.
- The method's reported robustness to noise, occlusion, and mislabeling suggests it could be deployed in real-world settings like medical robots, where clean data is rare.
- The fuzzy inference systems could be reused to construct new emotion datasets by applying the same rules to other facial expression databases, reducing annotation cost.
- The framework's efficiency (20.3 FPS on an edge device) points toward real-time emotion recognition on embedded platforms.
- The authors' claim of being 'first' to consider spatio-temporal heterogeneity in fine-grained emotion recognition sets a direction for future work that explicitly models changing emotion patterns over time and across scenarios.
Reading between the lines
- A strong internal consistency issue arises: the authors generate the ground-truth labels for DISFA-FER and WFLW-FER using the same fuzzy rules that the model uses for feature augmentation, so the model's success on those two benchmarks may partly reflect self-consistency rather than independent recognition of emotions.
- The fuzzy rules are hand-coded based on a fixed set of facial components and attribute values; whether these rules generalize to other modalities, cultures, or non-face-based emotional expressions is untested.
- A natural extension would be to learn the fuzzy membership functions and rules from data instead of fixing them; the paper's fitness-function tuning of λ1 and λ2 starts in that direction but does not learn the rules themselves.
- The robustness gains might partly come from the meta-learning task construction that groups multiple views of the same emotion, which could implicitly average out label noise, rather than from the fuzzy semantics alone; an ablation that keeps the meta-learning and spatio-temporal modules but removes only the fuzzy features would separate these contributions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ST-F2M, a meta-learning framework for fine-grained emotion recognition from multimodal video. Videos are segmented into modality-specific views with positional encoding, grouped into meta-learning tasks, encoded by a spatio-temporal convolutional module, augmented with fuzzy semantic information from two hand-built fuzzy inference systems (FCIS and FKIS), and optimized with a bi-level MAML-style procedure. The authors additionally construct two datasets, DISFA-FER and WFLW-FER, by applying FCIS/FKIS rules to existing face datasets. Experiments compare ST-F2M with a wide range of baselines on five datasets (CK+, DISFA-FER, WFLW-FER, CMU-MOSEI, CREMA-D), report robustness to noise, mislabeling, and few-shot settings, study computational efficiency on an edge platform, and include a medical-robot deployment case study.
Significance. If the empirical claims hold, the paper would make a useful contribution to fine-grained emotion recognition by demonstrating that a meta-learning formulation combined with fuzzy semantic priors can improve accuracy, robustness, and efficiency. The paper has several strengths: it formulates three concrete challenges (data annotation cost, temporal heterogeneity, spatial heterogeneity), provides a complete pipeline with pseudo-code, evaluates against a broad set of baselines, includes ablations for each module, and reports a real hardware deployment. The construction of two new fine-grained emotion datasets is also potentially valuable. However, the significance is currently limited because two of the five benchmark results are obtained on datasets whose labels were generated by the same fuzzy rules embedded in the model, making those results a form of self-consistency rather than independent recognition. The contribution would be substantially strengthened by independent validation of the constructed labels or by reframing the empirical claims around the three externally labeled datasets.
major comments (3)
- [§IV, §V-B3, §VI-A2, Table III]
- [§VI-F, Eq. (5)]
- [Table III, §VI-B]
minor comments (6)
- [§VI-A2]
- [Figures 2–3]
- [§VI-C]
- [§VI-F]
- [§V-B4]
- [References]
Circularity Check
DISFA-FER and WFLW-FER labels are generated by the same FCIS/FKIS rule base embedded in the model, making two of the five benchmark comparisons self-consistency checks rather than independent evaluations.
-
self definitional
[Section IV (second stated goal of FCIS/FKIS); Section VI-A2 (datasets); Section V-B3 (model inference path)]
"(ii) for dataset construction: assign emotion and intensity information to benchmark datasets in areas such as facial expression recognition to build benchmark datasets for fine-grained emotion recognition. / Note that for the facial expression datasets, DISFA and WFLW, we assign emotion and intensity information to the training data based on generalized fuzzy rules, and named the datasets DISFA-FER and WFLW-FER."
Section IV explicitly gives FCIS/FKIS two goals: 'for the decision-making process of ST-F2M' and 'for dataset construction: assign emotion and intensity information to benchmark datasets.' Section VI-A2 confirms DISFA-FER and WFLW-FER were made with the second use: 'we assign emotion and intensity information to the training data based on generalized fuzzy rules.' Section V-B3 shows the same systems inside the predictor: 'Zi is input into FCIS... Next, Zi is input into FKIS.' Hence on these two datasets, the ground-truth label is produced by the same rule tables embedded in the model.
full rationale
The clearest circular step is structural and documented by the paper itself. FCIS and FKIS are used both to build the DISFA-FER and WFLW-FER labels (Section VI-A2) and as a fixed stage of the ST-F2M prediction path (Section V-B3, Algorithm 1). On those two datasets, the evaluation therefore measures how well the model's learned encoder reproduces the facial-component coding assumed by the same hand-coded rule base that generated the labels. This makes the reported accuracy/recall advantages on DISFA-FER and WFLW-FER a self-consistency result, not an independent test of emotion recognition. The remaining benchmarks, CK+, CMU-MOSEI, and CREMA-D, use human labels and provide independent evidence for the method, so the circularity is partial rather than total. No load-bearing self-citation chain was found; the fuzzy-rule loop is the main issue. Overall score 7 reflects that two of the five headline benchmark results reduce by construction while the central method retains independent support on the other three datasets.
Assumptions & free parameters
free parameters (3)
- lambda_1 (FCIS membership range) =
0.4
- lambda_2 (FKIS membership range) =
0.4
- Membership function crossover points and rule values in FCIS and FKIS =
Hand-set values, e.g., eccentricity thresholds 0.15 and 0.45 for Angry in Fig. 3
assumptions (3)
- domain assumption Emotions can be represented as combinations of the 12 facial component attributes in Table I, each valued in {-1, 0, 1}.
- domain assumption The FKIS rule table (Table II) and the membership functions in Fig. 3 correctly map facial component sequences to the 18 emotion-intensity classes.
- domain assumption Bi-level meta-learning optimization with an inner loop on the support set and an outer loop on the query set transfers to fine-grained emotion recognition tasks.
Cite this review
Pith. "Pith review of Spatio-Temporal Fuzzy-oriented Multi-Modal Meta-Learning for Fine-grained Emotion Recognition." pith.science (2026). https://pith.science/paper/ZLF3BFQ7
@misc{pith2026241213541,
author = {Pith},
title = {Pith review of: Spatio-Temporal Fuzzy-oriented Multi-Modal Meta-Learning for Fine-grained Emotion Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZLF3BFQ7}},
note = {Machine review of arXiv:2412.13541}
}
read the original abstract
Fine-grained emotion recognition (FER) plays a vital role in various fields, such as disease diagnosis, personalized recommendations, and multimedia mining. However, existing FER methods face three key challenges in real-world applications: (i) they rely on large amounts of continuously annotated data to ensure accuracy since emotions are complex and ambiguous in reality, which is costly and time-consuming; (ii) they cannot capture the temporal heterogeneity caused by changing emotion patterns, because they usually assume that the temporal correlation within sampling periods is the same; (iii) they do not consider the spatial heterogeneity of different FER scenarios, that is, the distribution of emotion information in different data may have bias or interference. To address these challenges, we propose a Spatio-Temporal Fuzzy-oriented Multi-modal Meta-learning framework (ST-F2M). Specifically, ST-F2M first divides the multi-modal videos into multiple views, and each view corresponds to one modality of one emotion. Multiple randomly selected views for the same emotion form a meta-training task. Next, ST-F2M uses an integrated module with spatial and temporal convolutions to encode the data of each task, reflecting the spatial and temporal heterogeneity. Then it adds fuzzy semantic information to each task based on generalized fuzzy rules, which helps handle the complexity and ambiguity of emotions. Finally, ST-F2M learns emotion-related general meta-knowledge through meta-recurrent neural networks to achieve fast and robust fine-grained emotion recognition. Extensive experiments show that ST-F2M outperforms various state-of-the-art methods in terms of accuracy and model efficiency. In addition, we construct ablation studies and further analysis to explore why ST-F2M performs well.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Multi- modal emotion recognition with temporal and semantic consistency,
B. Chen, Q. Cao, M. Hou, Z. Zhang, G. Lu, and D. Zhang, “Multi- modal emotion recognition with temporal and semantic consistency,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3592–3603, 2021
2021
-
[2]
Graphcfc: A directed graph based cross-modal feature complementation approach for multimodal conversational emotion recognition,
J. Li, X. Wang, G. Lv, and Z. Zeng, “Graphcfc: A directed graph based cross-modal feature complementation approach for multimodal conversational emotion recognition,”IEEE Transactions on Multimedia, 2023
2023
-
[3]
Emocov: Machine learning for emotion detection, analysis and visualization using covid-19 tweets,
M. Y . Kabir and S. Madria, “Emocov: Machine learning for emotion detection, analysis and visualization using covid-19 tweets,”Online Social Networks and Media, vol. 23, p. 100135, 2021
2021
-
[4]
Recognition of emotions in user-generated videos through frame-level adaptation and emotion intensity learning,
H. Zhang and M. Xu, “Recognition of emotions in user-generated videos through frame-level adaptation and emotion intensity learning,” IEEE Transactions on Multimedia, vol. 25, pp. 881–891, 2021
2021
-
[5]
Customer preferences extraction for air purifiers based on fine-grained sentiment analysis of online reviews,
J. Zhang, A. Zhang, D. Liu, and Y . Bian, “Customer preferences extraction for air purifiers based on fine-grained sentiment analysis of online reviews,”Knowledge-Based Systems, vol. 228, p. 107259, 2021
2021
-
[6]
Self-organizing double function-link fuzzy brain emotional control system design for uncertain nonlinear systems,
T.-T. Huynh, C.-M. Lin, T.-L. Le, N.-Q.-K. Le, V .-P. Vu, and F. Chao, “Self-organizing double function-link fuzzy brain emotional control system design for uncertain nonlinear systems,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 52, no. 3, pp. 1852–1868, 2020
2020
-
[7]
Intelligent cockpit for intelligent vehicle in metaverse: A case study of empathetic auditory regulation of human emotion,
W. Li, L. Wu, C. Wang, J. Xue, W. Hu, S. Li, G. Guo, and D. Cao, “Intelligent cockpit for intelligent vehicle in metaverse: A case study of empathetic auditory regulation of human emotion,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 53, no. 4, pp. 2173– 2187, 2022
2022
-
[8]
Subtype-aware unsupervised domain adap- tation for medical diagnosis,
X. Liu, X. Liu, B. Hu, W. Ji, F. Xing, J. Lu, J. You, C.-C. J. Kuo, G. El Fakhri, and J. Woo, “Subtype-aware unsupervised domain adap- tation for medical diagnosis,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 3, 2021, pp. 2189–2197
2021
Show all 98 references
-
[9]
Fine-grained image analysis with deep learning: A survey,
X.-S. Wei, Y .-Z. Song, O. Mac Aodha, J. Wu, Y . Peng, J. Tang, J. Yang, and S. Belongie, “Fine-grained image analysis with deep learning: A survey,”IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 12, pp. 8927–8948, 2021
2021
-
[10]
Disturbance rejection in mimo systems with emotional-learning-based controller: Application to variable rotor-speed helicopters,
B. Debnath and S. Mija, “Disturbance rejection in mimo systems with emotional-learning-based controller: Application to variable rotor-speed helicopters,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 53, no. 7, pp. 4381–4392, 2023
2023
-
[11]
Brain network manifold learned by cognition-inspired graph embedding model for emotion recognition,
C. Li, P. Li, Z. Chen, L. Yang, F. Li, F. Wan, Z. Cao, D. Yao, B.-L. Lu, and P. Xu, “Brain network manifold learned by cognition-inspired graph embedding model for emotion recognition,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024. IEEE TRANSACTIONS ON SYST...
2024
-
[12]
Combining a parallel 2d cnn with a self-attention dilated residual network for ctc-based discrete speech emotion recognition,
Z. Zhao, Q. Li, Z. Zhang, N. Cummins, H. Wang, J. Tao, and B. W. Schuller, “Combining a parallel 2d cnn with a self-attention dilated residual network for ctc-based discrete speech emotion recognition,” Neural Networks, vol. 141, pp. 52–60, 2021
2021
-
[13]
Mlg-ncs: Multimodal local–global neuromorphic computing system for affective video content analysis,
X. Ji, Z. Dong, G. Zhou, C. S. Lai, and D. Qi, “Mlg-ncs: Multimodal local–global neuromorphic computing system for affective video content analysis,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024
2024
-
[14]
User preference mining based on fine-grained sentiment analysis,
Y . Xiao, C. Li, M. Th ¨urer, Y . Liu, and T. Qu, “User preference mining based on fine-grained sentiment analysis,”Journal of Retailing and Consumer Services, vol. 68, p. 103013, 2022
2022
-
[15]
Few-shot learning for fine-grained emotion recognition using physiological signals,
T. Zhang, A. El Ali, A. Hanjalic, and P. Cesar, “Few-shot learning for fine-grained emotion recognition using physiological signals,”IEEE Transactions on Multimedia, 2022
2022
-
[16]
Challenges in the diagnosis of parkinson’s disease,
E. Tolosa, A. Garrido, S. W. Scholz, and W. Poewe, “Challenges in the diagnosis of parkinson’s disease,”The Lancet Neurology, vol. 20, no. 5, pp. 385–397, 2021
2021
-
[17]
Multimodal emotion classification with multi-level semantic reasoning network,
T. Zhu, L. Li, J. Yang, S. Zhao, and X. Xiao, “Multimodal emotion classification with multi-level semantic reasoning network,”IEEE Transactions on Multimedia, 2022
2022
-
[18]
Unimse: Towards unified multimodal sentiment analysis and emotion recognition,
G. Hu, T.-E. Lin, Y . Zhao, G. Lu, Y . Wu, and Y . Li, “Unimse: Towards unified multimodal sentiment analysis and emotion recognition,”arXiv preprint arXiv:2211.11256, 2022
2022 arXiv
-
[19]
Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,
A. Zadeh, R. Zellers, E. Pincus, and L.-P. Morency, “Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,”arXiv preprint arXiv:1606.06259, 2016
2016 arXiv
-
[20]
Learning to augment expressions for few-shot fine- grained facial expression recognition,
W. Wang, Y . Fu, Q. Sun, T. Chen, C. Cao, Z. Zheng, G. Xu, H. Qiu, Y .-G. Jiang, and X. Xue, “Learning to augment expressions for few-shot fine- grained facial expression recognition,”arXiv preprint arXiv:2001.06144, 2020
2001 arXiv
-
[21]
Emotion recognition in conversation: Research challenges, datasets, and recent advances,
S. Poria, N. Majumder, R. Mihalcea, and E. Hovy, “Emotion recognition in conversation: Research challenges, datasets, and recent advances,” IEEE Access, vol. 7, pp. 100 943–100 953, 2019
2019
-
[22]
Feda: Fine- grained emotion difference analysis for facial expression recognition,
H. Liu, H. Cai, Q. Lin, X. Zhang, X. Li, and H. Xiao, “Feda: Fine- grained emotion difference analysis for facial expression recognition,” Biomedical Signal Processing and Control, vol. 79, p. 104209, 2023
2023
-
[23]
Fine-grained facial expression recognition in the wild,
L. Liang, C. Lang, Y . Li, S. Feng, and J. Zhao, “Fine-grained facial expression recognition in the wild,”IEEE Transactions on Information Forensics and Security, vol. 16, pp. 482–494, 2020
2020
-
[24]
Training deep networks for facial expression recognition with crowd-sourced label distribution,
E. Barsoum, C. Zhang, C. C. Ferrer, and Z. Zhang, “Training deep networks for facial expression recognition with crowd-sourced label distribution,” inProceedings of the 18th ACM International Conference on Multimodal Interaction, 2016, pp. 279–283
2016
-
[25]
Adaptive weighting of handcrafted feature losses for facial expression recognition,
W. Xie, L. Shen, and J. Duan, “Adaptive weighting of handcrafted feature losses for facial expression recognition,”IEEE transactions on cybernetics, vol. 51, no. 5, pp. 2787–2800, 2019
2019
-
[26]
Mixed feelings: expression of non-basic emotions in a muscle-based talking head,
I. Albrecht, M. Schr ¨oder, J. Haber, and H.-P. Seidel, “Mixed feelings: expression of non-basic emotions in a muscle-based talking head,” Virtual Reality, vol. 8, no. 4, pp. 201–212, 2005
2005
-
[27]
Convolutional features-based broad learning with lstm for multidimensional facial emotion recognition in human–robot interaction,
L. Chen, M. Li, M. Wu, W. Pedrycz, and K. Hirota, “Convolutional features-based broad learning with lstm for multidimensional facial emotion recognition in human–robot interaction,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 54, no. 1, pp. 64–75, 2023
2023
-
[28]
A review of affective computing: From unimodal analysis to multimodal fusion,
S. Poria, E. Cambria, R. Bajpai, and A. Hussain, “A review of affective computing: From unimodal analysis to multimodal fusion,”Information Fusion, vol. 37, pp. 98–125, 2017
2017
-
[29]
Design and analysis of a closed-loop emotion regulation system based on multimodal affective computing and emotional markov chain,
X. Wang, C.-Z. Li, Z. Sun, and Y . Xu, “Design and analysis of a closed-loop emotion regulation system based on multimodal affective computing and emotional markov chain,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2025
2025
-
[30]
Long dialogue emotion detection based on commonsense knowledge graph guidance,
W. Nie, Y . Bao, Y . Zhao, and A. Liu, “Long dialogue emotion detection based on commonsense knowledge graph guidance,”IEEE Transactions on Multimedia, 2023
2023
-
[31]
Emotional expression: Advances in basic emotion theory,
D. Keltner, D. Sauter, J. Tracy, and A. Cowen, “Emotional expression: Advances in basic emotion theory,”Journal of nonverbal behavior, vol. 43, no. 2, pp. 133–160, 2019
2019
-
[32]
Multimodal spontaneous emotion corpus for human behavior analysis,
Z. Zhang, J. M. Girard, Y . Wu, X. Zhang, P. Liu, U. Ciftci, S. Canavan, M. Reale, A. Horowitz, H. Yanget al., “Multimodal spontaneous emotion corpus for human behavior analysis,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 3438–3446
2016
-
[33]
A survey of textual emotion detection,
S. Al-Saqqa, H. Abdel-Nabi, and A. Awajan, “A survey of textual emotion detection,” in2018 8th International Conference on Computer Science and Information Technology (CSIT). IEEE, 2018, pp. 136–142
2018
-
[34]
Facial expression recognition based on deep learning,
H. Ge, Z. Zhu, Y . Dai, B. Wang, and X. Wu, “Facial expression recognition based on deep learning,”Computer Methods and Programs in Biomedicine, vol. 215, p. 106621, 2022
2022
-
[35]
A real time facial expression classification system using local binary patterns,
S. Happy, A. George, and A. Routray, “A real time facial expression classification system using local binary patterns,” in2012 4th Interna- tional conference on intelligent human computer interaction (IHCI). IEEE, 2012, pp. 1–5
2012
-
[36]
Human facial expression recognition using stepwise linear discriminant analysis and hidden conditional random fields,
M. H. Siddiqi, R. Ali, A. M. Khan, Y .-T. Park, and S. Lee, “Human facial expression recognition using stepwise linear discriminant analysis and hidden conditional random fields,”IEEE Transactions on Image Processing, vol. 24, no. 4, pp. 1386–1398, 2015
2015
-
[37]
Facial expression recognition based on local region specific features and support vector machines,
D. Ghimire, S. Jeong, J. Lee, and S. H. Park, “Facial expression recognition based on local region specific features and support vector machines,”Multimedia Tools and Applications, vol. 76, no. 6, pp. 7803– 7821, 2017
2017
-
[38]
A brief review of facial emotion recognition based on visual information,
B. C. Ko, “A brief review of facial emotion recognition based on visual information,”sensors, vol. 18, no. 2, p. 401, 2018
2018
-
[39]
Geometric-convolutional feature fusion based on learning propagation for facial expression recognition,
Y . Tang, X. M. Zhang, and H. Wang, “Geometric-convolutional feature fusion based on learning propagation for facial expression recognition,” IEEE Access, vol. 6, pp. 42 532–42 540, 2018
2018
-
[40]
Ga-svm-based facial emotion recognition using facial geometric features,
X. Liu, X. Cheng, and K. Lee, “Ga-svm-based facial emotion recognition using facial geometric features,”IEEE Sensors Journal, vol. 21, no. 10, pp. 11 532–11 542, 2020
2020
-
[41]
Graph based feature extraction and hybrid classification approach for facial expression recognition,
L. Krithika and G. Priya, “Graph based feature extraction and hybrid classification approach for facial expression recognition,”Journal of ambient intelligence and humanized computing, vol. 12, no. 2, pp. 2131–2147, 2021
2021
-
[42]
Mer-gcn: Micro- expression recognition based on relation modeling with graph convolu- tional networks,
L. Lo, H.-X. Xie, H.-H. Shuai, and W.-H. Cheng, “Mer-gcn: Micro- expression recognition based on relation modeling with graph convolu- tional networks,” in2020 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR). IEEE, 2020, pp. 79–84
2020
-
[43]
Facial expression recog- nition based on deep evolutional spatial-temporal networks,
K. Zhang, Y . Huang, Y . Du, and L. Wang, “Facial expression recog- nition based on deep evolutional spatial-temporal networks,”IEEE Transactions on Image Processing, vol. 26, no. 9, pp. 4193–4203, 2017
2017
-
[44]
Fg-agr: Fine- grained associative graph representation for facial expression recognition in the wild,
C. Li, X. Li, X. Wang, D. Huang, Z. Liu, and L. Liao, “Fg-agr: Fine- grained associative graph representation for facial expression recognition in the wild,”IEEE Transactions on Circuits and Systems for Video Technology, 2023
2023
-
[45]
Are there basic emotions?
P. Ekman, “Are there basic emotions?” 1992
1992
-
[46]
Eeg-based emotion recognition: A state-of-the-art review of current trends and opportunities,
N. S. Suhaimi, J. Mountstephens, J. Teoet al., “Eeg-based emotion recognition: A state-of-the-art review of current trends and opportunities,” Computational intelligence and neuroscience, vol. 2020, 2020
2020
-
[47]
Gut microbiota and fear processing in women affected by obesity: An exploratory pilot study,
F. Scarpina, S. Turroni, S. Mambrini, M. Barone, S. Cattaldo, S. Mai, E. Prina, I. Bastoni, S. Cappelli, G. Castelnuovoet al., “Gut microbiota and fear processing in women affected by obesity: An exploratory pilot study,”Nutrients, vol. 14, no. 18, p. 3788, 2022
2022
-
[48]
Emo-sensory commu- nication, emo-sensory intelligence and gender,
E. Naji Meidani, H. Makiabadi, M. Zabetipour, H. Abbasnejad, A. Firoozian Pooresfehani, and S. Shayesteh, “Emo-sensory commu- nication, emo-sensory intelligence and gender,”Journal of Business, Communication & Technology, vol. 1, no. 2, pp. 54–66, 2022
2022
-
[49]
Emotion recognition from geometric fuzzy member- ship functions,
R. Vishnu Priya, “Emotion recognition from geometric fuzzy member- ship functions,”Multimedia Tools and Applications, vol. 78, no. 13, pp. 17 847–17 878, 2019
2019
-
[50]
Group based emotion recognition from video sequence with hybrid optimization based recurrent fuzzy neural network,
V . Sreenivas, V . Namdeo, and E. V . Kumar, “Group based emotion recognition from video sequence with hybrid optimization based recurrent fuzzy neural network,”Journal of Big Data, vol. 7, no. 1, pp. 1–21, 2020
2020
-
[51]
Human emotion recognition based on active appearance model and semi-supervised fuzzy c-means,
D. Y . Liliana, M. R. Widyanto, and T. Basaruddin, “Human emotion recognition based on active appearance model and semi-supervised fuzzy c-means,” in2016 international conference on advanced computer science and information systems (ICACSIS). IEEE, 2016, pp. 439–445
2016
-
[52]
A review on machine learning styles in computer vision-techniques and future directions,
S. V . Mahadevkar, B. Khemani, S. Patil, K. Kotecha, D. V ora, A. Abraham, and L. A. Gabralla, “A review on machine learning styles in computer vision-techniques and future directions,”IEEE Access, 2022
2022
-
[53]
Model-agnostic meta-learning for fast adaptation of deep networks,
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” inInternational conference on machine learning. PMLR, 2017, pp. 1126–1135
2017
-
[55]
Rapid learning or feature reuse? towards understanding the effectiveness of maml,
A. Raghu, M. Raghu, S. Bengio, and O. Vinyals, “Rapid learning or feature reuse? towards understanding the effectiveness of maml,”arXiv preprint arXiv:1909.09157, 2019
1909 arXiv
-
[56]
Meta-learning with memory-augmented neural networks,
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap, “Meta-learning with memory-augmented neural networks,” inInterna- tional conference on machine learning. PMLR, 2016, pp. 1842–1850. IEEE TRANSACTIONS ON SYSTEMS, MAN AND CYBERNETICS: SYSTEMS 13
2016
-
[57]
Meta-learning improves lifelong rela- tion extraction,
A. Obamuyide and A. Vlachos, “Meta-learning improves lifelong rela- tion extraction,” inProceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), 2019, pp. 224–229
2019
-
[60]
Variational metric scaling for metric-based meta-learning,
J. Chen, L.-M. Zhan, X.-M. Wu, and F.-l. Chung, “Variational metric scaling for metric-based meta-learning,” inProceedings of the AAAI conference on artificial intelligence, vol. 34, no. 04, 2020, pp. 3478– 3485
2020
-
[61]
Deep neural network for emotion recognition based on meta-transfer learning,
H. Tang, G. Jiang, and Q. Wang, “Deep neural network for emotion recognition based on meta-transfer learning,”IEEE Access, vol. 10, pp. 78 114–78 122, 2022
2022
-
[62]
Emotion recognition from facial images with simultaneous occlusion, pose and illumination variations using meta-learning,
S. Kuruvayil and S. Palaniswamy, “Emotion recognition from facial images with simultaneous occlusion, pose and illumination variations using meta-learning,”Journal of King Saud University-Computer and Information Sciences, vol. 34, no. 9, pp. 7271–7282, 2022
2022
-
[63]
Meta-transfer learning for emotion recogni- tion,
D. Nguyen, D. T. Nguyen, S. Sridharan, S. Denman, T. T. Nguyen, D. Dean, and C. Fookes, “Meta-transfer learning for emotion recogni- tion,”Neural Computing and Applications, pp. 1–15, 2023
2023
-
[64]
Cluster-level contrastive learning for emotion recognition in conversations,
K. Yang, T. Zhang, H. Alhuzali, and S. Ananiadou, “Cluster-level contrastive learning for emotion recognition in conversations,”IEEE Transactions on Affective Computing, 2023
2023
-
[65]
The psychology of emotion regulation: An integrative review,
S. L. Koole, “The psychology of emotion regulation: An integrative review,”Cognition and emotion, vol. 23, no. 1, pp. 4–41, 2009
2009
-
[66]
Comparison of mamdani-type and sugeno-type fuzzy inference systems for air conditioning system,
A. Kaur and A. Kaur, “Comparison of mamdani-type and sugeno-type fuzzy inference systems for air conditioning system,”International Journal of Soft Computing and Engineering (IJSCE), vol. 2, no. 2, pp. 323–325, 2012
2012
-
[67]
Disfa: A spontaneous facial action intensity database,
S. M. Mavadati, M. H. Mahoor, K. Bartlett, P. Trinh, and J. F. Cohn, “Disfa: A spontaneous facial action intensity database,”IEEE Transactions on Affective Computing, vol. 4, no. 2, pp. 151–160, 2013
2013
-
[68]
Look at boundary: A boundary-aware face alignment algorithm,
W. Wu, C. Qian, S. Yang, Q. Wang, Y . Cai, and Q. Zhou, “Look at boundary: A boundary-aware face alignment algorithm,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 2129–2138
2018
-
[69]
Spatio-temporal graph convolutional networks: A deep learning framework for traffic forecasting,
B. Yu, H. Yin, and Z. Zhu, “Spatio-temporal graph convolutional networks: A deep learning framework for traffic forecasting,”arXiv preprint arXiv:1709.04875, 2017
2017 arXiv
-
[70]
Rethinking graph neural architecture search from message-passing,
S. Cai, L. Li, J. Deng, B. Zhang, Z.-J. Zha, L. Su, and Q. Huang, “Rethinking graph neural architecture search from message-passing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 6657–6666
2021
-
[71]
Fitness functions in evolutionary robotics: A survey and analysis,
A. L. Nelson, G. J. Barlow, and L. Doitsidis, “Fitness functions in evolutionary robotics: A survey and analysis,”Robotics and Autonomous Systems, vol. 57, no. 4, pp. 345–370, 2009
2009
-
[72]
Fitness function design to improve evolutionary structural testing,
A. Baresel, H. Sthamer, and M. Schmidt, “Fitness function design to improve evolutionary structural testing,” inProceedings of the 4th Annual Conference on Genetic and Evolutionary Computation, 2002, pp. 1329–1336
2002
-
[73]
Improving adam optimizer,
A. Tato and R. Nkambou, “Improving adam optimizer,” 2018
2018
-
[74]
The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,
P. Lucey, J. F. Cohn, T. Kanade, J. Saragih, Z. Ambadar, and I. Matthews, “The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,” in2010 ieee computer society conference on computer vision and pattern recognition-workshops...
2010
-
[75]
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,
A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” inProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa...
2018
-
[76]
Crema-d: Crowd-sourced emotional multimodal actors dataset,
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,”IEEE transactions on affective computing, vol. 5, no. 4, pp. 377–390, 2014
2014
-
[77]
Tailor versatile multi-modal learning for multi-label emotion recognition,
Y . Zhang, M. Chen, J. Shen, and C. Wang, “Tailor versatile multi-modal learning for multi-label emotion recognition,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 8, 2022, pp. 9100–9108
2022
-
[78]
Openface: an open source facial behavior analysis toolkit,
T. Baltruˇsaitis, P. Robinson, and L.-P. Morency, “Openface: an open source facial behavior analysis toolkit,” in2016 IEEE winter conference on applications of computer vision (WACV). IEEE, 2016, pp. 1–10
2016
-
[79]
Glove: Global vectors for word representation,
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” inProceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014, pp. 1532–1543
2014
-
[80]
Facenet2expnet: Regularizing a deep face recognition net for expression recognition,
H. Ding, S. K. Zhou, and R. Chellappa, “Facenet2expnet: Regularizing a deep face recognition net for expression recognition,” in2017 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017). IEEE, 2017, pp. 118–126
2017
-
[81]
Pre-training strategies and datasets for facial representation learning,
A. Bulat, S. Cheng, J. Yang, A. Garbett, E. Sanchez, and G. Tzimiropou- los, “Pre-training strategies and datasets for facial representation learning,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 107–125
2022
-
[82]
Propagationnet: Propagate points to curve to learn structure information,
X. Huang, W. Deng, H. Shen, X. Zhang, and J. Ye, “Propagationnet: Propagate points to curve to learn structure information,” inProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 7265–7274
2020
-
[83]
Septr: Separa- ble transformer for audio spectrogram processing,
N.-C. Ristea, R. T. Ionescu, and F. S. Khan, “Septr: Separa- ble transformer for audio spectrogram processing,”arXiv preprint arXiv:2203.09581, 2022
2022 arXiv
-
[84]
Reptile: a scalable metalearning algorithm,
A. Nichol and J. Schulman, “Reptile: a scalable metalearning algorithm,” arXiv preprint arXiv:1803.02999, vol. 2, no. 3, p. 4, 2018
2018 arXiv
-
[85]
Prototypical networks for few-shot learning,
J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,”Advances in neural information processing systems, vol. 30, 2017
2017
-
[86]
Learning to compare: Relation network for few-shot learning,
F. Sung, Y . Yang, L. Zhang, T. Xiang, P. H. Torr, and T. M. Hospedales, “Learning to compare: Relation network for few-shot learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 1199–1208
2018
-
[87]
Hacking task confounder in meta-learning,
J. Wang, W. Qiang, Y . Ren, Z. Song, X. Su, and C. Zheng, “Hacking task confounder in meta-learning,”arXiv preprint arXiv:2312.05771, 2023
2023 arXiv
-
[88]
Combining deep and unsupervised features for multilingual speech emotion recognition,
V . Scotti, F. Galati, L. Sbattella, and R. Tedesco, “Combining deep and unsupervised features for multilingual speech emotion recognition,” in International Conference on Pattern Recognition. Springer, 2021, pp. 114–128
2021
-
[89]
Cogmen: Contextualized gnn based multimodal emotion recognition,
A. Joshi, A. Bhat, A. Jain, A. V . Singh, and A. Modi, “Cogmen: Contextualized gnn based multimodal emotion recognition,”arXiv preprint arXiv:2205.02455, 2022
2022 arXiv
-
[90]
Marlin: Masked autoencoder for facial video representation learning,
Z. Cai, S. Ghosh, K. Stefanov, A. Dhall, J. Cai, H. Rezatofighi, R. Haffari, and M. Hayat, “Marlin: Masked autoencoder for facial video representation learning,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1493–1504
2023
-
[91]
Emotion recognition based on brain-like multimodal hierarchical perception,
X. Zhu, Y . Huang, X. Wang, and R. Wang, “Emotion recognition based on brain-like multimodal hierarchical perception,”Multimedia Tools and Applications, vol. 83, no. 18, pp. 56 039–56 057, 2024
2024
-
[92]
Fusing pairwise modalities for emotion recognition in conversations,
C. Fan, J. Lin, R. Mao, and E. Cambria, “Fusing pairwise modalities for emotion recognition in conversations,”Information Fusion, vol. 106, p. 102306, 2024
2024
-
[93]
A novel transformer autoencoder for multi-modal emotion recognition with incomplete data,
C. Cheng, W. Liu, Z. Fan, L. Feng, and Z. Jia, “A novel transformer autoencoder for multi-modal emotion recognition with incomplete data,” Neural Networks, vol. 172, p. 106111, 2024
2024
-
[94]
Deep imbalanced learning for multimodal emotion recognition in conversations,
T. Meng, Y . Shou, W. Ai, N. Yin, and K. Li, “Deep imbalanced learning for multimodal emotion recognition in conversations,”IEEE Transactions on Artificial Intelligence, 2024
2024
-
[95]
Online bagging and boosting,
N. C. Oza and S. J. Russell, “Online bagging and boosting,” in International Workshop on Artificial Intelligence and Statistics. PMLR, 2001, pp. 229–236
2001
-
[96]
Appli- cations of machine learning predictive models in the chronic disease diagnosis,
G. Battineni, G. G. Sagaro, N. Chinatalapudi, and F. Amenta, “Appli- cations of machine learning predictive models in the chronic disease diagnosis,”Journal of personalized medicine, vol. 10, no. 2, p. 21, 2020
2020
-
[97]
Effective image enhancement techniques for fog-affected indoor and outdoor images,
K. Kim, S. Kim, and K.-S. Kim, “Effective image enhancement techniques for fog-affected indoor and outdoor images,”IET Image Processing, vol. 12, no. 4, pp. 465–471, 2018
2018
-
[98]
A new application of fractional atangana– baleanu derivatives: designing abc-fractional masks in image processing,
B. Ghanbari and A. Atangana, “A new application of fractional atangana– baleanu derivatives: designing abc-fractional masks in image processing,” Physica A: Statistical Mechanics and its Applications, vol. 542, p. 123516, 2020
2020
-
[99]
Distortion ag- nostic deep watermarking,
X. Luo, R. Zhan, H. Chang, F. Yang, and P. Milanfar, “Distortion ag- nostic deep watermarking,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 13 548–13 557
2020
-
[100]
A precision analysis of camera distortion models,
Z. Tang, R. G. V on Gioi, P. Monasse, and J.-M. Morel, “A precision analysis of camera distortion models,”IEEE Transactions on Image Processing, vol. 26, no. 6, pp. 2694–2704, 2017
2017
-
[101]
Occlusion- aware real-time object tracking,
X. Dong, J. Shen, D. Yu, W. Wang, J. Liu, and H. Huang, “Occlusion- aware real-time object tracking,”IEEE Transactions on Multimedia, vol. 19, no. 4, pp. 763–771, 2016
2016
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.