Pith. sign in

REVIEW 3 major objections 6 minor 98 references

Spatio-Temporal Fuzzy-oriented Multi-Modal Meta-Learning for Fine-grained Emotion Recognition

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A meta-learning framework with fuzzy rules and spatio-temporal encoding claims top accuracy for fine-grained emotion recognition, even with limited or noisy data.

desk verdict A genuinely integrated method whose two self-labeled benchmarks make the headline five-dataset claim weaker than it looks; the three external benchmarks still show real gains. read the letter →

arxiv 2412.13541 v4 pith:ZLF3BFQ7 submitted 2024-12-18 cs.CV cs.LGcs.NE

classification cs.CVcs.LGcs.NE
keywords fine-grainedemotionrecognitionmeta-learningfuzzyinferencespatio-temporalheterogeneitymulti-modallearningfacialexpressionintensity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes ST-F2M, a framework that combines spatio-temporal encoding, generalized fuzzy inference, and meta-learning to recognize fine-grained emotions (six basic emotions with three intensity levels) from multi-modal videos. It aims to overcome three real-world obstacles: the need for large amounts of annotated data, the assumption of constant temporal correlations, and the neglect of spatial differences across scenarios. The authors claim that ST-F2M beats state-of-the-art methods on five benchmarks while being faster and more robust to noise, mislabeling, and data scarcity.

What carries the argument

The key objects are the two fuzzy inference systems: the fuzzy component inference system (FCIS), which converts facial component states into linguistic condition sequences using membership functions, and the fuzzy knowledge inference system (FKIS), which maps those sequences to emotional intensity levels (low, medium, high) for each of six basic emotions. These fuzzy systems serve both as a feature augmentation module inside the model and as a rule-based label generator for two new benchmarks, DISFA-FER and WFLW-FER. The learning machinery is a meta-recurrent neural network with a spatio-temporal convolutional encoder that performs bi-level optimization over tasks constructed from multi-modal views.

What would settle it

Take a new facial expression dataset with human-annotated fine-grained emotion and intensity labels (not derived from the paper's fuzzy rules), train ST-F2M with only the feature-augmentation role of the fuzzy systems, and compare against a version with the fuzzy modules removed. If the fuzzy-augmented model does not outperform the ablated version on this independently labeled dataset, the claim that fuzzy semantics improve emotion recognition would be falsified.

Watch

Extended reading notes

Core claim

The central claim is that fine-grained emotion recognition can be made data-efficient and robust by decomposing emotions into fuzzy components, encoding spatio-temporal heterogeneity with convolutions, and learning general emotion meta-knowledge via bi-level optimization. The paper introduces two fuzzy inference systems (FCIS and FKIS) that quantify emotional intensity and assign fuzzy semantic information to representations, and it uses a meta-recurrent neural network to achieve rapid adaptation. On five datasets, ST-F2M reports accuracy improvements of up to several points over previous methods, and it maintains high performance under synthetic noise and with only 20% of training data plus 30% mislabeled samples.

Load-bearing premise

The hand-coded fuzzy rules in FCIS and FKIS are assumed to be a valid compositional model of emotion, and the ground-truth labels for two of the five benchmark datasets are generated from those same rules, so the strong results on those datasets could reflect self-consistency rather than independent recognition.

Editorial extensions

If this is right

  • If the claimed results hold, ST-F2M would enable fine-grained emotion recognition with far fewer annotated samples, since the fuzzy rules supply emotional intensity information without manual labeling.
  • The method's reported robustness to noise, occlusion, and mislabeling suggests it could be deployed in real-world settings like medical robots, where clean data is rare.
  • The fuzzy inference systems could be reused to construct new emotion datasets by applying the same rules to other facial expression databases, reducing annotation cost.
  • The framework's efficiency (20.3 FPS on an edge device) points toward real-time emotion recognition on embedded platforms.
  • The authors' claim of being 'first' to consider spatio-temporal heterogeneity in fine-grained emotion recognition sets a direction for future work that explicitly models changing emotion patterns over time and across scenarios.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A strong internal consistency issue arises: the authors generate the ground-truth labels for DISFA-FER and WFLW-FER using the same fuzzy rules that the model uses for feature augmentation, so the model's success on those two benchmarks may partly reflect self-consistency rather than independent recognition of emotions.
  • The fuzzy rules are hand-coded based on a fixed set of facial components and attribute values; whether these rules generalize to other modalities, cultures, or non-face-based emotional expressions is untested.
  • A natural extension would be to learn the fuzzy membership functions and rules from data instead of fixing them; the paper's fitness-function tuning of λ1 and λ2 starts in that direction but does not learn the rules themselves.
  • The robustness gains might partly come from the meta-learning task construction that groups multiple views of the same emotion, which could implicitly average out label noise, rather than from the fuzzy semantics alone; an ablation that keeps the meta-learning and spatio-temporal modules but removes only the fuzzy features would separate these contributions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes ST-F2M, a meta-learning framework for fine-grained emotion recognition from multimodal video. Videos are segmented into modality-specific views with positional encoding, grouped into meta-learning tasks, encoded by a spatio-temporal convolutional module, augmented with fuzzy semantic information from two hand-built fuzzy inference systems (FCIS and FKIS), and optimized with a bi-level MAML-style procedure. The authors additionally construct two datasets, DISFA-FER and WFLW-FER, by applying FCIS/FKIS rules to existing face datasets. Experiments compare ST-F2M with a wide range of baselines on five datasets (CK+, DISFA-FER, WFLW-FER, CMU-MOSEI, CREMA-D), report robustness to noise, mislabeling, and few-shot settings, study computational efficiency on an edge platform, and include a medical-robot deployment case study.

Significance. If the empirical claims hold, the paper would make a useful contribution to fine-grained emotion recognition by demonstrating that a meta-learning formulation combined with fuzzy semantic priors can improve accuracy, robustness, and efficiency. The paper has several strengths: it formulates three concrete challenges (data annotation cost, temporal heterogeneity, spatial heterogeneity), provides a complete pipeline with pseudo-code, evaluates against a broad set of baselines, includes ablations for each module, and reports a real hardware deployment. The construction of two new fine-grained emotion datasets is also potentially valuable. However, the significance is currently limited because two of the five benchmark results are obtained on datasets whose labels were generated by the same fuzzy rules embedded in the model, making those results a form of self-consistency rather than independent recognition. The contribution would be substantially strengthened by independent validation of the constructed labels or by reframing the empirical claims around the three externally labeled datasets.

major comments (3)
  1. [§IV, §V-B3, §VI-A2, Table III]
  2. [§VI-F, Eq. (5)]
  3. [Table III, §VI-B]
minor comments (6)
  1. [§VI-A2]
  2. [Figures 2–3]
  3. [§VI-C]
  4. [§VI-F]
  5. [§V-B4]
  6. [References]

Circularity Check

1 steps flagged · score 7.0 of 10

DISFA-FER and WFLW-FER labels are generated by the same FCIS/FKIS rule base embedded in the model, making two of the five benchmark comparisons self-consistency checks rather than independent evaluations.

  1. self definitional [Section IV (second stated goal of FCIS/FKIS); Section VI-A2 (datasets); Section V-B3 (model inference path)]
    "(ii) for dataset construction: assign emotion and intensity information to benchmark datasets in areas such as facial expression recognition to build benchmark datasets for fine-grained emotion recognition. / Note that for the facial expression datasets, DISFA and WFLW, we assign emotion and intensity information to the training data based on generalized fuzzy rules, and named the datasets DISFA-FER and WFLW-FER."

    Section IV explicitly gives FCIS/FKIS two goals: 'for the decision-making process of ST-F2M' and 'for dataset construction: assign emotion and intensity information to benchmark datasets.' Section VI-A2 confirms DISFA-FER and WFLW-FER were made with the second use: 'we assign emotion and intensity information to the training data based on generalized fuzzy rules.' Section V-B3 shows the same systems inside the predictor: 'Zi is input into FCIS... Next, Zi is input into FKIS.' Hence on these two datasets, the ground-truth label is produced by the same rule tables embedded in the model.

full rationale

The clearest circular step is structural and documented by the paper itself. FCIS and FKIS are used both to build the DISFA-FER and WFLW-FER labels (Section VI-A2) and as a fixed stage of the ST-F2M prediction path (Section V-B3, Algorithm 1). On those two datasets, the evaluation therefore measures how well the model's learned encoder reproduces the facial-component coding assumed by the same hand-coded rule base that generated the labels. This makes the reported accuracy/recall advantages on DISFA-FER and WFLW-FER a self-consistency result, not an independent test of emotion recognition. The remaining benchmarks, CK+, CMU-MOSEI, and CREMA-D, use human labels and provide independent evidence for the method, so the circularity is partial rather than total. No load-bearing self-citation chain was found; the fuzzy-rule loop is the main issue. Overall score 7 reflects that two of the five headline benchmark results reduce by construction while the central method retains independent support on the other three datasets.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The framework relies on a set of hand-designed fuzzy rules and membership functions. Two of these are treated as free parameters (lambda_1 and lambda_2), and the membership function geometry itself is chosen by hand. The domain assumptions are that emotions decompose into facial component states and that the FKIS rule table maps those states to emotional intensity. No new physical entities are introduced.

free parameters (3)
  • lambda_1 (FCIS membership range) = 0.4
    Selected by parameter investigation on the five benchmarks in Section VI-F; used in Eq. 5 to weight FCIS membership functions.
  • lambda_2 (FKIS membership range) = 0.4
    Same tuning procedure as lambda_1; the best joint setting in the reported 3D histogram is 0.4 for both.
  • Membership function crossover points and rule values in FCIS and FKIS = Hand-set values, e.g., eccentricity thresholds 0.15 and 0.45 for Angry in Fig. 3
    Chosen by hand from psychological knowledge in Section IV and Figs. 2-3; these values directly determine the fuzzy encodings and also the labels of DISFA-FER and WFLW-FER.
assumptions (3)
  • domain assumption Emotions can be represented as combinations of the 12 facial component attributes in Table I, each valued in {-1, 0, 1}.
    Invoked in Section IV and Table I; the whole FCIS system and the dataset labeling depend on this decomposition.
  • domain assumption The FKIS rule table (Table II) and the membership functions in Fig. 3 correctly map facial component sequences to the 18 emotion-intensity classes.
    Invoked in Section IV; used both as model features and as ground-truth labels for DISFA-FER and WFLW-FER.
  • domain assumption Bi-level meta-learning optimization with an inner loop on the support set and an outer loop on the query set transfers to fine-grained emotion recognition tasks.
    Adopted in Section III and Eq. 6-7 from MAML-style methods; assumed rather than demonstrated by comparison to alternative learning paradigms.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Spatio-Temporal Fuzzy-oriented Multi-Modal Meta-Learning for Fine-grained Emotion Recognition." pith.science (2026). https://pith.science/paper/ZLF3BFQ7

@misc{pith2026241213541,
  author       = {Pith},
  title        = {Pith review of: Spatio-Temporal Fuzzy-oriented Multi-Modal Meta-Learning for Fine-grained Emotion Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZLF3BFQ7}},
  note         = {Machine review of arXiv:2412.13541}
}
read the original abstract

Fine-grained emotion recognition (FER) plays a vital role in various fields, such as disease diagnosis, personalized recommendations, and multimedia mining. However, existing FER methods face three key challenges in real-world applications: (i) they rely on large amounts of continuously annotated data to ensure accuracy since emotions are complex and ambiguous in reality, which is costly and time-consuming; (ii) they cannot capture the temporal heterogeneity caused by changing emotion patterns, because they usually assume that the temporal correlation within sampling periods is the same; (iii) they do not consider the spatial heterogeneity of different FER scenarios, that is, the distribution of emotion information in different data may have bias or interference. To address these challenges, we propose a Spatio-Temporal Fuzzy-oriented Multi-modal Meta-learning framework (ST-F2M). Specifically, ST-F2M first divides the multi-modal videos into multiple views, and each view corresponds to one modality of one emotion. Multiple randomly selected views for the same emotion form a meta-training task. Next, ST-F2M uses an integrated module with spatial and temporal convolutions to encode the data of each task, reflecting the spatial and temporal heterogeneity. Then it adds fuzzy semantic information to each task based on generalized fuzzy rules, which helps handle the complexity and ambiguity of emotions. Finally, ST-F2M learns emotion-related general meta-knowledge through meta-recurrent neural networks to achieve fast and robust fine-grained emotion recognition. Extensive experiments show that ST-F2M outperforms various state-of-the-art methods in terms of accuracy and model efficiency. In addition, we construct ablation studies and further analysis to explore why ST-F2M performs well.

Figures

Figures reproduced from arXiv: 2412.13541 by the authors.

Figure 1
Figure 1. Illustration of our motivation, i.e., the three key challenges in fine [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The membership functions for all the physiological attributes. The horizontal axis represents the input, such as the feature encoding result of the neural [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. The membership functions example for the intensity value of six emotions. The horizontal axis represents the converted eccentricity value of the [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The framework of ST-F2M. ST-F2M first constructs multi-modal meta-learning tasks based on the input videos. Next, it performs encoding through a [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Trade-off performance. It provides the comparative results of models’ [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: The confusion matrix of six emotions [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 10
Figure 10. Figure 10: Parameter investigations about {λ1, λ2}. (a) Data collection Group A Group B Group C Group D Group E 40 50 60 70 80 90 100 Accuracy FN2EN UniMSE SepTr COGMEN MARLIN MAML ANIL ST-F2M (b) Comparison results [PITH_FULL_IMAGE:figures/full_fig_p009_10.png]
Figure 11
Figure 11. Figure 11: Comparison results of practical case study. Group A [PITH_FULL_IMAGE:figures/full_fig_p009_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

98 extracted references · 64 canonical work pages

  1. [1]

    Multi- modal emotion recognition with temporal and semantic consistency,

    B. Chen, Q. Cao, M. Hou, Z. Zhang, G. Lu, and D. Zhang, “Multi- modal emotion recognition with temporal and semantic consistency,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3592–3603, 2021

  2. [2]

    Graphcfc: A directed graph based cross-modal feature complementation approach for multimodal conversational emotion recognition,

    J. Li, X. Wang, G. Lv, and Z. Zeng, “Graphcfc: A directed graph based cross-modal feature complementation approach for multimodal conversational emotion recognition,”IEEE Transactions on Multimedia, 2023

  3. [3]

    Emocov: Machine learning for emotion detection, analysis and visualization using covid-19 tweets,

    M. Y . Kabir and S. Madria, “Emocov: Machine learning for emotion detection, analysis and visualization using covid-19 tweets,”Online Social Networks and Media, vol. 23, p. 100135, 2021

  4. [4]

    Recognition of emotions in user-generated videos through frame-level adaptation and emotion intensity learning,

    H. Zhang and M. Xu, “Recognition of emotions in user-generated videos through frame-level adaptation and emotion intensity learning,” IEEE Transactions on Multimedia, vol. 25, pp. 881–891, 2021

  5. [5]

    Customer preferences extraction for air purifiers based on fine-grained sentiment analysis of online reviews,

    J. Zhang, A. Zhang, D. Liu, and Y . Bian, “Customer preferences extraction for air purifiers based on fine-grained sentiment analysis of online reviews,”Knowledge-Based Systems, vol. 228, p. 107259, 2021

  6. [6]

    Self-organizing double function-link fuzzy brain emotional control system design for uncertain nonlinear systems,

    T.-T. Huynh, C.-M. Lin, T.-L. Le, N.-Q.-K. Le, V .-P. Vu, and F. Chao, “Self-organizing double function-link fuzzy brain emotional control system design for uncertain nonlinear systems,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 52, no. 3, pp. 1852–1868, 2020

  7. [7]

    Intelligent cockpit for intelligent vehicle in metaverse: A case study of empathetic auditory regulation of human emotion,

    W. Li, L. Wu, C. Wang, J. Xue, W. Hu, S. Li, G. Guo, and D. Cao, “Intelligent cockpit for intelligent vehicle in metaverse: A case study of empathetic auditory regulation of human emotion,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 53, no. 4, pp. 2173– 2187, 2022

  8. [8]

    Subtype-aware unsupervised domain adap- tation for medical diagnosis,

    X. Liu, X. Liu, B. Hu, W. Ji, F. Xing, J. Lu, J. You, C.-C. J. Kuo, G. El Fakhri, and J. Woo, “Subtype-aware unsupervised domain adap- tation for medical diagnosis,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 3, 2021, pp. 2189–2197

Show all 98 references
  1. [9]

    Fine-grained image analysis with deep learning: A survey,

    X.-S. Wei, Y .-Z. Song, O. Mac Aodha, J. Wu, Y . Peng, J. Tang, J. Yang, and S. Belongie, “Fine-grained image analysis with deep learning: A survey,”IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 12, pp. 8927–8948, 2021

  2. [10]

    Disturbance rejection in mimo systems with emotional-learning-based controller: Application to variable rotor-speed helicopters,

    B. Debnath and S. Mija, “Disturbance rejection in mimo systems with emotional-learning-based controller: Application to variable rotor-speed helicopters,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 53, no. 7, pp. 4381–4392, 2023

  3. [11]

    Brain network manifold learned by cognition-inspired graph embedding model for emotion recognition,

    C. Li, P. Li, Z. Chen, L. Yang, F. Li, F. Wan, Z. Cao, D. Yao, B.-L. Lu, and P. Xu, “Brain network manifold learned by cognition-inspired graph embedding model for emotion recognition,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024. IEEE TRANSACTIONS ON SYST...

  4. [12]

    Combining a parallel 2d cnn with a self-attention dilated residual network for ctc-based discrete speech emotion recognition,

    Z. Zhao, Q. Li, Z. Zhang, N. Cummins, H. Wang, J. Tao, and B. W. Schuller, “Combining a parallel 2d cnn with a self-attention dilated residual network for ctc-based discrete speech emotion recognition,” Neural Networks, vol. 141, pp. 52–60, 2021

  5. [13]

    Mlg-ncs: Multimodal local–global neuromorphic computing system for affective video content analysis,

    X. Ji, Z. Dong, G. Zhou, C. S. Lai, and D. Qi, “Mlg-ncs: Multimodal local–global neuromorphic computing system for affective video content analysis,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024

  6. [14]

    User preference mining based on fine-grained sentiment analysis,

    Y . Xiao, C. Li, M. Th ¨urer, Y . Liu, and T. Qu, “User preference mining based on fine-grained sentiment analysis,”Journal of Retailing and Consumer Services, vol. 68, p. 103013, 2022

  7. [15]

    Few-shot learning for fine-grained emotion recognition using physiological signals,

    T. Zhang, A. El Ali, A. Hanjalic, and P. Cesar, “Few-shot learning for fine-grained emotion recognition using physiological signals,”IEEE Transactions on Multimedia, 2022

  8. [16]

    Challenges in the diagnosis of parkinson’s disease,

    E. Tolosa, A. Garrido, S. W. Scholz, and W. Poewe, “Challenges in the diagnosis of parkinson’s disease,”The Lancet Neurology, vol. 20, no. 5, pp. 385–397, 2021

  9. [17]

    Multimodal emotion classification with multi-level semantic reasoning network,

    T. Zhu, L. Li, J. Yang, S. Zhao, and X. Xiao, “Multimodal emotion classification with multi-level semantic reasoning network,”IEEE Transactions on Multimedia, 2022

  10. [18]

    Unimse: Towards unified multimodal sentiment analysis and emotion recognition,

    G. Hu, T.-E. Lin, Y . Zhao, G. Lu, Y . Wu, and Y . Li, “Unimse: Towards unified multimodal sentiment analysis and emotion recognition,”arXiv preprint arXiv:2211.11256, 2022

  11. [19]

    Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,

    A. Zadeh, R. Zellers, E. Pincus, and L.-P. Morency, “Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,”arXiv preprint arXiv:1606.06259, 2016

  12. [20]

    Learning to augment expressions for few-shot fine- grained facial expression recognition,

    W. Wang, Y . Fu, Q. Sun, T. Chen, C. Cao, Z. Zheng, G. Xu, H. Qiu, Y .-G. Jiang, and X. Xue, “Learning to augment expressions for few-shot fine- grained facial expression recognition,”arXiv preprint arXiv:2001.06144, 2020

  13. [21]

    Emotion recognition in conversation: Research challenges, datasets, and recent advances,

    S. Poria, N. Majumder, R. Mihalcea, and E. Hovy, “Emotion recognition in conversation: Research challenges, datasets, and recent advances,” IEEE Access, vol. 7, pp. 100 943–100 953, 2019

  14. [22]

    Feda: Fine- grained emotion difference analysis for facial expression recognition,

    H. Liu, H. Cai, Q. Lin, X. Zhang, X. Li, and H. Xiao, “Feda: Fine- grained emotion difference analysis for facial expression recognition,” Biomedical Signal Processing and Control, vol. 79, p. 104209, 2023

  15. [23]

    Fine-grained facial expression recognition in the wild,

    L. Liang, C. Lang, Y . Li, S. Feng, and J. Zhao, “Fine-grained facial expression recognition in the wild,”IEEE Transactions on Information Forensics and Security, vol. 16, pp. 482–494, 2020

  16. [24]

    Training deep networks for facial expression recognition with crowd-sourced label distribution,

    E. Barsoum, C. Zhang, C. C. Ferrer, and Z. Zhang, “Training deep networks for facial expression recognition with crowd-sourced label distribution,” inProceedings of the 18th ACM International Conference on Multimodal Interaction, 2016, pp. 279–283

  17. [25]

    Adaptive weighting of handcrafted feature losses for facial expression recognition,

    W. Xie, L. Shen, and J. Duan, “Adaptive weighting of handcrafted feature losses for facial expression recognition,”IEEE transactions on cybernetics, vol. 51, no. 5, pp. 2787–2800, 2019

  18. [26]

    Mixed feelings: expression of non-basic emotions in a muscle-based talking head,

    I. Albrecht, M. Schr ¨oder, J. Haber, and H.-P. Seidel, “Mixed feelings: expression of non-basic emotions in a muscle-based talking head,” Virtual Reality, vol. 8, no. 4, pp. 201–212, 2005

  19. [27]

    Convolutional features-based broad learning with lstm for multidimensional facial emotion recognition in human–robot interaction,

    L. Chen, M. Li, M. Wu, W. Pedrycz, and K. Hirota, “Convolutional features-based broad learning with lstm for multidimensional facial emotion recognition in human–robot interaction,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 54, no. 1, pp. 64–75, 2023

  20. [28]

    A review of affective computing: From unimodal analysis to multimodal fusion,

    S. Poria, E. Cambria, R. Bajpai, and A. Hussain, “A review of affective computing: From unimodal analysis to multimodal fusion,”Information Fusion, vol. 37, pp. 98–125, 2017

  21. [29]

    Design and analysis of a closed-loop emotion regulation system based on multimodal affective computing and emotional markov chain,

    X. Wang, C.-Z. Li, Z. Sun, and Y . Xu, “Design and analysis of a closed-loop emotion regulation system based on multimodal affective computing and emotional markov chain,”IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2025

  22. [30]

    Long dialogue emotion detection based on commonsense knowledge graph guidance,

    W. Nie, Y . Bao, Y . Zhao, and A. Liu, “Long dialogue emotion detection based on commonsense knowledge graph guidance,”IEEE Transactions on Multimedia, 2023

  23. [31]

    Emotional expression: Advances in basic emotion theory,

    D. Keltner, D. Sauter, J. Tracy, and A. Cowen, “Emotional expression: Advances in basic emotion theory,”Journal of nonverbal behavior, vol. 43, no. 2, pp. 133–160, 2019

  24. [32]

    Multimodal spontaneous emotion corpus for human behavior analysis,

    Z. Zhang, J. M. Girard, Y . Wu, X. Zhang, P. Liu, U. Ciftci, S. Canavan, M. Reale, A. Horowitz, H. Yanget al., “Multimodal spontaneous emotion corpus for human behavior analysis,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 3438–3446

  25. [33]

    A survey of textual emotion detection,

    S. Al-Saqqa, H. Abdel-Nabi, and A. Awajan, “A survey of textual emotion detection,” in2018 8th International Conference on Computer Science and Information Technology (CSIT). IEEE, 2018, pp. 136–142

  26. [34]

    Facial expression recognition based on deep learning,

    H. Ge, Z. Zhu, Y . Dai, B. Wang, and X. Wu, “Facial expression recognition based on deep learning,”Computer Methods and Programs in Biomedicine, vol. 215, p. 106621, 2022

  27. [35]

    A real time facial expression classification system using local binary patterns,

    S. Happy, A. George, and A. Routray, “A real time facial expression classification system using local binary patterns,” in2012 4th Interna- tional conference on intelligent human computer interaction (IHCI). IEEE, 2012, pp. 1–5

  28. [36]

    Human facial expression recognition using stepwise linear discriminant analysis and hidden conditional random fields,

    M. H. Siddiqi, R. Ali, A. M. Khan, Y .-T. Park, and S. Lee, “Human facial expression recognition using stepwise linear discriminant analysis and hidden conditional random fields,”IEEE Transactions on Image Processing, vol. 24, no. 4, pp. 1386–1398, 2015

  29. [37]

    Facial expression recognition based on local region specific features and support vector machines,

    D. Ghimire, S. Jeong, J. Lee, and S. H. Park, “Facial expression recognition based on local region specific features and support vector machines,”Multimedia Tools and Applications, vol. 76, no. 6, pp. 7803– 7821, 2017

  30. [38]

    A brief review of facial emotion recognition based on visual information,

    B. C. Ko, “A brief review of facial emotion recognition based on visual information,”sensors, vol. 18, no. 2, p. 401, 2018

  31. [39]

    Geometric-convolutional feature fusion based on learning propagation for facial expression recognition,

    Y . Tang, X. M. Zhang, and H. Wang, “Geometric-convolutional feature fusion based on learning propagation for facial expression recognition,” IEEE Access, vol. 6, pp. 42 532–42 540, 2018

  32. [40]

    Ga-svm-based facial emotion recognition using facial geometric features,

    X. Liu, X. Cheng, and K. Lee, “Ga-svm-based facial emotion recognition using facial geometric features,”IEEE Sensors Journal, vol. 21, no. 10, pp. 11 532–11 542, 2020

  33. [41]

    Graph based feature extraction and hybrid classification approach for facial expression recognition,

    L. Krithika and G. Priya, “Graph based feature extraction and hybrid classification approach for facial expression recognition,”Journal of ambient intelligence and humanized computing, vol. 12, no. 2, pp. 2131–2147, 2021

  34. [42]

    Mer-gcn: Micro- expression recognition based on relation modeling with graph convolu- tional networks,

    L. Lo, H.-X. Xie, H.-H. Shuai, and W.-H. Cheng, “Mer-gcn: Micro- expression recognition based on relation modeling with graph convolu- tional networks,” in2020 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR). IEEE, 2020, pp. 79–84

  35. [43]

    Facial expression recog- nition based on deep evolutional spatial-temporal networks,

    K. Zhang, Y . Huang, Y . Du, and L. Wang, “Facial expression recog- nition based on deep evolutional spatial-temporal networks,”IEEE Transactions on Image Processing, vol. 26, no. 9, pp. 4193–4203, 2017

  36. [44]

    Fg-agr: Fine- grained associative graph representation for facial expression recognition in the wild,

    C. Li, X. Li, X. Wang, D. Huang, Z. Liu, and L. Liao, “Fg-agr: Fine- grained associative graph representation for facial expression recognition in the wild,”IEEE Transactions on Circuits and Systems for Video Technology, 2023

  37. [45]

    Are there basic emotions?

    P. Ekman, “Are there basic emotions?” 1992

  38. [46]

    Eeg-based emotion recognition: A state-of-the-art review of current trends and opportunities,

    N. S. Suhaimi, J. Mountstephens, J. Teoet al., “Eeg-based emotion recognition: A state-of-the-art review of current trends and opportunities,” Computational intelligence and neuroscience, vol. 2020, 2020

  39. [47]

    Gut microbiota and fear processing in women affected by obesity: An exploratory pilot study,

    F. Scarpina, S. Turroni, S. Mambrini, M. Barone, S. Cattaldo, S. Mai, E. Prina, I. Bastoni, S. Cappelli, G. Castelnuovoet al., “Gut microbiota and fear processing in women affected by obesity: An exploratory pilot study,”Nutrients, vol. 14, no. 18, p. 3788, 2022

  40. [48]

    Emo-sensory commu- nication, emo-sensory intelligence and gender,

    E. Naji Meidani, H. Makiabadi, M. Zabetipour, H. Abbasnejad, A. Firoozian Pooresfehani, and S. Shayesteh, “Emo-sensory commu- nication, emo-sensory intelligence and gender,”Journal of Business, Communication & Technology, vol. 1, no. 2, pp. 54–66, 2022

  41. [49]

    Emotion recognition from geometric fuzzy member- ship functions,

    R. Vishnu Priya, “Emotion recognition from geometric fuzzy member- ship functions,”Multimedia Tools and Applications, vol. 78, no. 13, pp. 17 847–17 878, 2019

  42. [50]

    Group based emotion recognition from video sequence with hybrid optimization based recurrent fuzzy neural network,

    V . Sreenivas, V . Namdeo, and E. V . Kumar, “Group based emotion recognition from video sequence with hybrid optimization based recurrent fuzzy neural network,”Journal of Big Data, vol. 7, no. 1, pp. 1–21, 2020

  43. [51]

    Human emotion recognition based on active appearance model and semi-supervised fuzzy c-means,

    D. Y . Liliana, M. R. Widyanto, and T. Basaruddin, “Human emotion recognition based on active appearance model and semi-supervised fuzzy c-means,” in2016 international conference on advanced computer science and information systems (ICACSIS). IEEE, 2016, pp. 439–445

  44. [52]

    A review on machine learning styles in computer vision-techniques and future directions,

    S. V . Mahadevkar, B. Khemani, S. Patil, K. Kotecha, D. V ora, A. Abraham, and L. A. Gabralla, “A review on machine learning styles in computer vision-techniques and future directions,”IEEE Access, 2022

  45. [53]

    Model-agnostic meta-learning for fast adaptation of deep networks,

    C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” inInternational conference on machine learning. PMLR, 2017, pp. 1126–1135

  46. [55]

    Rapid learning or feature reuse? towards understanding the effectiveness of maml,

    A. Raghu, M. Raghu, S. Bengio, and O. Vinyals, “Rapid learning or feature reuse? towards understanding the effectiveness of maml,”arXiv preprint arXiv:1909.09157, 2019

  47. [56]

    Meta-learning with memory-augmented neural networks,

    A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap, “Meta-learning with memory-augmented neural networks,” inInterna- tional conference on machine learning. PMLR, 2016, pp. 1842–1850. IEEE TRANSACTIONS ON SYSTEMS, MAN AND CYBERNETICS: SYSTEMS 13

  48. [57]

    Meta-learning improves lifelong rela- tion extraction,

    A. Obamuyide and A. Vlachos, “Meta-learning improves lifelong rela- tion extraction,” inProceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), 2019, pp. 224–229

  49. [60]

    Variational metric scaling for metric-based meta-learning,

    J. Chen, L.-M. Zhan, X.-M. Wu, and F.-l. Chung, “Variational metric scaling for metric-based meta-learning,” inProceedings of the AAAI conference on artificial intelligence, vol. 34, no. 04, 2020, pp. 3478– 3485

  50. [61]

    Deep neural network for emotion recognition based on meta-transfer learning,

    H. Tang, G. Jiang, and Q. Wang, “Deep neural network for emotion recognition based on meta-transfer learning,”IEEE Access, vol. 10, pp. 78 114–78 122, 2022

  51. [62]

    Emotion recognition from facial images with simultaneous occlusion, pose and illumination variations using meta-learning,

    S. Kuruvayil and S. Palaniswamy, “Emotion recognition from facial images with simultaneous occlusion, pose and illumination variations using meta-learning,”Journal of King Saud University-Computer and Information Sciences, vol. 34, no. 9, pp. 7271–7282, 2022

  52. [63]

    Meta-transfer learning for emotion recogni- tion,

    D. Nguyen, D. T. Nguyen, S. Sridharan, S. Denman, T. T. Nguyen, D. Dean, and C. Fookes, “Meta-transfer learning for emotion recogni- tion,”Neural Computing and Applications, pp. 1–15, 2023

  53. [64]

    Cluster-level contrastive learning for emotion recognition in conversations,

    K. Yang, T. Zhang, H. Alhuzali, and S. Ananiadou, “Cluster-level contrastive learning for emotion recognition in conversations,”IEEE Transactions on Affective Computing, 2023

  54. [65]

    The psychology of emotion regulation: An integrative review,

    S. L. Koole, “The psychology of emotion regulation: An integrative review,”Cognition and emotion, vol. 23, no. 1, pp. 4–41, 2009

  55. [66]

    Comparison of mamdani-type and sugeno-type fuzzy inference systems for air conditioning system,

    A. Kaur and A. Kaur, “Comparison of mamdani-type and sugeno-type fuzzy inference systems for air conditioning system,”International Journal of Soft Computing and Engineering (IJSCE), vol. 2, no. 2, pp. 323–325, 2012

  56. [67]

    Disfa: A spontaneous facial action intensity database,

    S. M. Mavadati, M. H. Mahoor, K. Bartlett, P. Trinh, and J. F. Cohn, “Disfa: A spontaneous facial action intensity database,”IEEE Transactions on Affective Computing, vol. 4, no. 2, pp. 151–160, 2013

  57. [68]

    Look at boundary: A boundary-aware face alignment algorithm,

    W. Wu, C. Qian, S. Yang, Q. Wang, Y . Cai, and Q. Zhou, “Look at boundary: A boundary-aware face alignment algorithm,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 2129–2138

  58. [69]

    Spatio-temporal graph convolutional networks: A deep learning framework for traffic forecasting,

    B. Yu, H. Yin, and Z. Zhu, “Spatio-temporal graph convolutional networks: A deep learning framework for traffic forecasting,”arXiv preprint arXiv:1709.04875, 2017

  59. [70]

    Rethinking graph neural architecture search from message-passing,

    S. Cai, L. Li, J. Deng, B. Zhang, Z.-J. Zha, L. Su, and Q. Huang, “Rethinking graph neural architecture search from message-passing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 6657–6666

  60. [71]

    Fitness functions in evolutionary robotics: A survey and analysis,

    A. L. Nelson, G. J. Barlow, and L. Doitsidis, “Fitness functions in evolutionary robotics: A survey and analysis,”Robotics and Autonomous Systems, vol. 57, no. 4, pp. 345–370, 2009

  61. [72]

    Fitness function design to improve evolutionary structural testing,

    A. Baresel, H. Sthamer, and M. Schmidt, “Fitness function design to improve evolutionary structural testing,” inProceedings of the 4th Annual Conference on Genetic and Evolutionary Computation, 2002, pp. 1329–1336

  62. [73]

    Improving adam optimizer,

    A. Tato and R. Nkambou, “Improving adam optimizer,” 2018

  63. [74]

    The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,

    P. Lucey, J. F. Cohn, T. Kanade, J. Saragih, Z. Ambadar, and I. Matthews, “The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,” in2010 ieee computer society conference on computer vision and pattern recognition-workshops...

  64. [75]

    Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,

    A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” inProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa...

  65. [76]

    Crema-d: Crowd-sourced emotional multimodal actors dataset,

    H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,”IEEE transactions on affective computing, vol. 5, no. 4, pp. 377–390, 2014

  66. [77]

    Tailor versatile multi-modal learning for multi-label emotion recognition,

    Y . Zhang, M. Chen, J. Shen, and C. Wang, “Tailor versatile multi-modal learning for multi-label emotion recognition,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 8, 2022, pp. 9100–9108

  67. [78]

    Openface: an open source facial behavior analysis toolkit,

    T. Baltruˇsaitis, P. Robinson, and L.-P. Morency, “Openface: an open source facial behavior analysis toolkit,” in2016 IEEE winter conference on applications of computer vision (WACV). IEEE, 2016, pp. 1–10

  68. [79]

    Glove: Global vectors for word representation,

    J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” inProceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014, pp. 1532–1543

  69. [80]

    Facenet2expnet: Regularizing a deep face recognition net for expression recognition,

    H. Ding, S. K. Zhou, and R. Chellappa, “Facenet2expnet: Regularizing a deep face recognition net for expression recognition,” in2017 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017). IEEE, 2017, pp. 118–126

  70. [81]

    Pre-training strategies and datasets for facial representation learning,

    A. Bulat, S. Cheng, J. Yang, A. Garbett, E. Sanchez, and G. Tzimiropou- los, “Pre-training strategies and datasets for facial representation learning,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 107–125

  71. [82]

    Propagationnet: Propagate points to curve to learn structure information,

    X. Huang, W. Deng, H. Shen, X. Zhang, and J. Ye, “Propagationnet: Propagate points to curve to learn structure information,” inProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 7265–7274

  72. [83]

    Septr: Separa- ble transformer for audio spectrogram processing,

    N.-C. Ristea, R. T. Ionescu, and F. S. Khan, “Septr: Separa- ble transformer for audio spectrogram processing,”arXiv preprint arXiv:2203.09581, 2022

  73. [84]

    Reptile: a scalable metalearning algorithm,

    A. Nichol and J. Schulman, “Reptile: a scalable metalearning algorithm,” arXiv preprint arXiv:1803.02999, vol. 2, no. 3, p. 4, 2018

  74. [85]

    Prototypical networks for few-shot learning,

    J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,”Advances in neural information processing systems, vol. 30, 2017

  75. [86]

    Learning to compare: Relation network for few-shot learning,

    F. Sung, Y . Yang, L. Zhang, T. Xiang, P. H. Torr, and T. M. Hospedales, “Learning to compare: Relation network for few-shot learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 1199–1208

  76. [87]

    Hacking task confounder in meta-learning,

    J. Wang, W. Qiang, Y . Ren, Z. Song, X. Su, and C. Zheng, “Hacking task confounder in meta-learning,”arXiv preprint arXiv:2312.05771, 2023

  77. [88]

    Combining deep and unsupervised features for multilingual speech emotion recognition,

    V . Scotti, F. Galati, L. Sbattella, and R. Tedesco, “Combining deep and unsupervised features for multilingual speech emotion recognition,” in International Conference on Pattern Recognition. Springer, 2021, pp. 114–128

  78. [89]

    Cogmen: Contextualized gnn based multimodal emotion recognition,

    A. Joshi, A. Bhat, A. Jain, A. V . Singh, and A. Modi, “Cogmen: Contextualized gnn based multimodal emotion recognition,”arXiv preprint arXiv:2205.02455, 2022

  79. [90]

    Marlin: Masked autoencoder for facial video representation learning,

    Z. Cai, S. Ghosh, K. Stefanov, A. Dhall, J. Cai, H. Rezatofighi, R. Haffari, and M. Hayat, “Marlin: Masked autoencoder for facial video representation learning,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1493–1504

  80. [91]

    Emotion recognition based on brain-like multimodal hierarchical perception,

    X. Zhu, Y . Huang, X. Wang, and R. Wang, “Emotion recognition based on brain-like multimodal hierarchical perception,”Multimedia Tools and Applications, vol. 83, no. 18, pp. 56 039–56 057, 2024

  81. [92]

    Fusing pairwise modalities for emotion recognition in conversations,

    C. Fan, J. Lin, R. Mao, and E. Cambria, “Fusing pairwise modalities for emotion recognition in conversations,”Information Fusion, vol. 106, p. 102306, 2024

  82. [93]

    A novel transformer autoencoder for multi-modal emotion recognition with incomplete data,

    C. Cheng, W. Liu, Z. Fan, L. Feng, and Z. Jia, “A novel transformer autoencoder for multi-modal emotion recognition with incomplete data,” Neural Networks, vol. 172, p. 106111, 2024

  83. [94]

    Deep imbalanced learning for multimodal emotion recognition in conversations,

    T. Meng, Y . Shou, W. Ai, N. Yin, and K. Li, “Deep imbalanced learning for multimodal emotion recognition in conversations,”IEEE Transactions on Artificial Intelligence, 2024

  84. [95]

    Online bagging and boosting,

    N. C. Oza and S. J. Russell, “Online bagging and boosting,” in International Workshop on Artificial Intelligence and Statistics. PMLR, 2001, pp. 229–236

  85. [96]

    Appli- cations of machine learning predictive models in the chronic disease diagnosis,

    G. Battineni, G. G. Sagaro, N. Chinatalapudi, and F. Amenta, “Appli- cations of machine learning predictive models in the chronic disease diagnosis,”Journal of personalized medicine, vol. 10, no. 2, p. 21, 2020

  86. [97]

    Effective image enhancement techniques for fog-affected indoor and outdoor images,

    K. Kim, S. Kim, and K.-S. Kim, “Effective image enhancement techniques for fog-affected indoor and outdoor images,”IET Image Processing, vol. 12, no. 4, pp. 465–471, 2018

  87. [98]

    A new application of fractional atangana– baleanu derivatives: designing abc-fractional masks in image processing,

    B. Ghanbari and A. Atangana, “A new application of fractional atangana– baleanu derivatives: designing abc-fractional masks in image processing,” Physica A: Statistical Mechanics and its Applications, vol. 542, p. 123516, 2020

  88. [99]

    Distortion ag- nostic deep watermarking,

    X. Luo, R. Zhan, H. Chang, F. Yang, and P. Milanfar, “Distortion ag- nostic deep watermarking,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 13 548–13 557

  89. [100]

    A precision analysis of camera distortion models,

    Z. Tang, R. G. V on Gioi, P. Monasse, and J.-M. Morel, “A precision analysis of camera distortion models,”IEEE Transactions on Image Processing, vol. 26, no. 6, pp. 2694–2704, 2017

  90. [101]

    Occlusion- aware real-time object tracking,

    X. Dong, J. Shen, D. Yu, W. Wang, J. Liu, and H. Huang, “Occlusion- aware real-time object tracking,”IEEE Transactions on Multimedia, vol. 19, no. 4, pp. 763–771, 2016

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.