REVIEW 4 major objections 5 minor 83 references
Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A prior-guided video-level regression network sets new state-of-the-art results in micro-expression analysis.
desk verdict Adaptive interval selection is a real improvement over fixed-window decoding, but undisclosed thresholds make the SOTA claim conditional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The scalable interval selection strategy (SISS): starting from a peak in the per-frame spotting probability, the interval is extended bidirectionally; within the prior duration range k a lower threshold prunes frames only after two consecutive below-threshold frames, while beyond k a higher threshold terminates the interval after two consecutive failures. It is what converts per-frame regression outputs into onset-apex-offset intervals without preset window lengths. The second mechanism is synergistic optimization: one SlowFast Mamba encoder shared by spotting and recognition, with separate FC heads.
What would settle it
Take the released code, run it on CAS(ME)^3, and sweep both thresholds over a grid; if the reported STRS of 0.0562 is only achieved for a narrow range of threshold pairs and collapses otherwise, the central claim of a robust prior-guided strategy is disconfirmed. Alternatively, a reproduction that uses only the described defaults (no per-dataset threshold tuning) and fails to reach the reported scores would falsify the claim that the method as described is state of the art.
Extended reading notes
Core claim
The central claim is that replacing window-based interval decoding with a prior-guided, threshold-driven interval selection, and letting spotting and recognition share every layer except their prediction heads, improves joint ME analysis. Concretely, the scalable interval selection strategy uses a dataset-derived average duration k and two thresholds: inside k a lower threshold keeps a frame as long as no two consecutive frames fall below it, and outside k a higher threshold is applied, so boundaries can extend flexibly. The paper reports STRS of 0.0562 on CAS(ME)^3 and 0.2000 on SAMMLV, outperforming prior simultaneous methods such as ME-TST, MEAN, and SFAMNet, and topping the MEGC2025 test
Load-bearing premise
The whole decoding scheme assumes that a dataset-derived average duration range k and two fixed probability threshold values, which are not reported, generalize across videos; if those numbers are tuned per benchmark, the 'prior-guided' advantage may not transfer.
Editorial extensions
If this is right
- Reported results imply that per-frame soft prediction plus threshold-based decoding can replace sliding-window classifiers entirely.
- Joint training with shared parameters means spotting and recognition reinforce each other, which is especially valuable with scarce micro-expression data.
- If the approach generalizes, long-video ME analysis systems could be built with a single lightweight network instead of separate stages.
- Improved IoU_tp and IoU_all suggest the strategy also corrects boundary localization, not just detection.
- The method needs fewer parameters than separate-task pipelines, so the same architecture can scale to longer videos at lower cost.
Reading between the lines
- If the two thresholds and the duration range k are stable across datasets, the same interval-selection rule could transfer to macro-expression spotting or other event-detection tasks in long videos.
- The two-level threshold mechanism could itself be learned from data, making the 'prior-guided' design fully adaptive instead of hand-set.
- A systematic sensitivity analysis of the thresholds would clarify whether the reported gains come from the regression framework or from tuning those hidden parameters.
- Because the paper reports the thinnest gains on the harder CAS(ME)^3 dataset, testing on a third long-video dataset with different frame rates would show whether the approach generalizes beyond the two benchmarks used for tuning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a video-level regression framework for simultaneous micro-expression spotting and recognition. It builds on the authors' earlier ME-TST+/ME-TST architecture (Section 3.1) and contributes two main ideas: (i) a Scalable Interval Selection Strategy (SISS) that decodes onset/apex/offset intervals from per-frame spotting probabilities using a two-level thresholding rule instead of a fixed window, and (ii) a Synergistic Optimization (SO) scheme that shares all network parameters between spotting and recognition except the final classification heads. The method is evaluated on CAS(ME)^3, SAMMLV, and the MEGC2025 test set under leave-one-subject-out and leaderboard protocols, reporting state-of-the-art STRS of 0.0562 on CAS(ME)^3 and 0.2000 on SAMMLV, and first place on the MEGC2025-testSet leaderboard.
Significance. If the reported results are robust, the paper would make a useful contribution to integrated micro-expression analysis: it replaces fixed-window decoding with an adaptive interval-selection rule and demonstrates that aggressive parameter sharing between spotting and recognition is beneficial in a low-data regime. The MEGC2025 leaderboard result provides some external validation beyond the authors' own evaluations, and the promise of released code is a positive step for reproducibility. However, the contribution's magnitude is currently difficult to assess because the key parameters of SISS are not disclosed, the backbone is taken from the authors' own later work, and no sensitivity analysis or statistical significance testing is provided. These issues limit the certainty with which the state-of-the-art claims can be accepted.
major comments (4)
- [Section 3.2, Table 5] The central contribution SISS is not fully specified. The average ME duration range k, the lower and higher probability thresholds, and the 'two consecutive frames' termination rule are never given numerical values, nor is it stated whether they are fixed across datasets, derived from training labels, or tuned on the test videos. Table 5 shows that SISS is the main driver of the CAS(ME)^3 improvement (STRS 0.0480 to 0.0562), and the same two-frame rule has very different temporal meaning at 30 fps versus 200 fps. Without these values and a sensitivity analysis, a substantial part of the reported gain may be an artifact of post-processing parameters rather than the framework itself. Please report all SISS parameters, explain how they are set, and show performance as a function of k and the thresholds.
- [Section 3.1, Table 3] The method follows the design of the authors' own ME-TST+ [16], but Table 3 compares against ME-TST [79] rather than ME-TST+. The reported STRS gains over ME-TST therefore conflate two factors: the upgraded backbone inherited from [16] (ROI relationship awareness, SlowFast Mamba) and the proposed SISS/SO contributions. To support the claim that SISS/SO are responsible for the improvement, Table 5 should include an ME-TST+ baseline without SISS and SO as a separate row, or the paper should clearly separate backbone contribution from post-processing/training contributions.
- [Table 4] The MEGC2025-testSet results do not consistently mirror the full-dataset margins. On the SAMMLV unseen subset, the proposed method obtains spotting F1=0.09, below USTC-IAT's 0.12, and the analysis STRS is tied at 0.06. This weakens the external-validity argument for the state-of-the-art claim and suggests that the full-dataset LOSO improvements may be dataset-specific or affected by protocol differences. Please discuss this discrepancy and report the SISS threshold behavior on the unseen subset.
- [Section 3.3, Section 4.2] Several training hyperparameters that can materially affect performance are not reported: the non-ME sampling ratio, the emotion-category loss weights, and the probability penalization factor used at inference. Additionally, no error bars or significance tests are provided for any table. Given the small absolute margins for some metrics (e.g., spotting F1 on CAS(ME)^3: 0.0997 vs 0.0802), the robustness of the conclusions is unclear. Please provide the missing hyperparameter values and add variance estimates or significance tests across LOSO folds.
minor comments (5)
- [Section 3.2] Minor typos and inconsistencies: 'se+lection' in the first paragraph; the text refers to 'average ME duration range k' but k is not defined formally. Please define all notation.
- [Section 4.1, Table 4] The column header 'Overall Analysis Analysis Spotting Recognition' in Table 4 appears to have a duplicated 'Analysis'; please correct the table formatting.
- [Section 4.3] The evaluation protocol is described only by references [22,29]. Please specify the exact LOSO splitting rule and the number of folds for each dataset to allow exact reproduction.
- [References [16] and [79]] The relationship between ME-TST+ [16] and ME-TST [79] should be stated explicitly in the methodology section, since the current text can leave the reader uncertain about which prior work the method extends.
- [Figure 2] The visualization provides qualitative support but the font sizes and axis labels are very small; please enlarge them so the onset/apex/offset comparisons are legible.
Circularity Check
No demonstrated circularity: the SOTA claim is empirically validated against external benchmarks; SISS parameters are undisclosed but not shown to be fitted to the target metric.
full rationale
I find no step in the paper where a claimed prediction reduces by construction to an input, fit, or self-citation. The central claim is empirical: the method is evaluated on CAS(ME)^3, SAMMLV, and the MEGC2025 test set, and the reported STRS values (e.g., 0.0562 and 0.2000 in Table 3) are computed from spotting and recognition F1 scores against ground-truth annotations, not from the method's own definitions. The Scalable Interval Selection Strategy (SISS) in Sec. 3.2 uses a duration prior k and two probability thresholds, but the paper does not state how these are obtained, whether they are tuned per dataset, or any sensitivity analysis. Missing parameter disclosure is a reproducibility weakness, not circularity: nothing in the text equates the thresholds with the test outputs or shows that the target metric is forced by a fitted constant. The backbone is inherited from the authors' own ME-TST+/ME-TST works (refs [16,79]), and the paper cites these for architecture and preprocessing, but it does not invoke a uniqueness theorem or use the self-citations to rule out alternatives. The contribution of the paper, SISS and synergistic optimization, is isolated by the ablation in Table 5 and the external leaderboard in Table 4, so the self-citations are not load-bearing for the novel components. Because the paper borrows heavily from the same group's prior architecture but the central SOTA claim rests on independent benchmark evaluation, the appropriate score is low; the omitted SISS parameter values and lack of sensitivity analysis keep it from being a clean 0, but no circular step is exhibited.
Assumptions & free parameters
free parameters (5)
- SISS duration range k =
not reported (average ME duration)
- SISS lower threshold =
not reported
- SISS higher threshold =
not reported
- Probability penalization factor =
not reported
- Recognition loss weights / non-ME sampling ratio =
not reported
assumptions (4)
- domain assumption Optical flow sequences from facial ROIs at multiple granularities carry enough signal for both spotting and recognition.
- domain assumption The average ME duration range k is a stable prior that transfers across datasets and can serve as the threshold switch point.
- domain assumption Spotting and recognition are complementary enough that hard parameter sharing, except for the heads, improves both tasks.
- domain assumption Per-frame spotting probabilities are sufficiently well-calibrated for fixed dual-threshold decoding.
Cite this review
Pith. "Pith review of Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression." pith.science (2026). https://pith.science/paper/2ZV3ZBV7
@misc{pith2026250818834,
author = {Pith},
title = {Pith review of: Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression},
year = {2026},
howpublished = {\url{https://pith.science/paper/2ZV3ZBV7}},
note = {Machine review of arXiv:2508.18834}
}
abstract
Micro-expressions (MEs) are involuntary, low-intensity, and short-duration facial expressions that often reveal an individual's genuine thoughts and emotions. Most existing ME analysis methods rely on window-level classification with fixed window sizes and hard decisions, which limits their ability to capture the complex temporal dynamics of MEs. Although recent approaches have adopted video-level regression frameworks to address some of these challenges, interval decoding still depends on manually predefined, window-based methods, leaving the issue only partially mitigated. In this paper, we propose a prior-guided video-level regression method for ME analysis. We introduce a scalable interval selection strategy that comprehensively considers the temporal evolution, duration, and class distribution characteristics of MEs, enabling precise spotting of the onset, apex, and offset phases. In addition, we introduce a synergistic optimization framework, in which the spotting and recognition tasks share parameters except for the classification heads. This fully exploits complementary information, makes more efficient use of limited data, and enhances the model's capability. Extensive experiments on multiple benchmark datasets demonstrate the state-of-the-art performance of our method, with an STRS of 0.0562 on CAS(ME)$^3$ and 0.2000 on SAMMLV. The code is available at https://github.com/zizheng-guo/BoostingVRME.
Figures
Reference graph
Works this paper leans on
-
[16]
Zizheng Guo, Bochao Zou, Junbao Zhuo, and Huimin Ma. 2025. ME-TST+: Micro-expression Analysis via Temporal State Transition with ROI Relationship Awareness. arXiv preprint arXiv:2508.08082 (2025)
arXiv 2025
-
[79]
Guoying Zhao and Matti Pietikainen. 2007. Dynamic texture recognition using local binary patterns with an application to facial expressions. IEEE transactions on pattern analysis and machine intelligence 29, 6 (2007), 915–928
work page 2007
-
[1]
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 6836–6846
2021
-
[2]
Xianye Ben, Yi Ren, Junping Zhang, Su-Jing Wang, Kidiyo Kpalma, Weixiao Meng, and Yong-Jin Liu. 2021. Video-based facial micro-expression analysis: A survey of datasets, features and algorithms. IEEE transactions on pattern analysis and machine intelligence 44, 9 (2021), 5826–5846
2021
-
[3]
Rizwan Chaudhry, Avinash Ravichandran, Gregory Hager, and René Vidal. 2009. Histograms of oriented optical flow and binet-cauchy kernels on nonlinear dy- namical systems for the recognition of human actions. In 2009 IEEE conference on computer vision and pattern recognition. IEEE, 1932–1939
2009
-
[4]
Tri Dao and Albert Gu. 2024. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. In International Conference on Machine Learning (ICML)
2024
-
[5]
Davison, Cliff Lansley, Nicholas Costen, Kevin Tan, and Moi Hoon Yap
Adrian K. Davison, Cliff Lansley, Nicholas Costen, Kevin Tan, and Moi Hoon Yap
-
[6]
Adrian K. Davison, Jingting Li, Moi Hoon Yap, John See, Wen-Huang Cheng, Xiaobai Li, Xiaopeng Hong, and Su-Jing Wang. 2023. MEGC2023: ACM Multi- media 2023 ME Grand Challenge. In Proceedings of the 31st ACM International Conference on Multimedia (Ottawa ON, Canada) (MM ’23). Association for Com- puting Machinery, New York, NY, USA, 9625–9629. doi:10.1145/...
Show all 83 references
-
[7]
Adrian K Davison, Moi Hoon Yap, and Cliff Lansley. 2015. Micro-facial movement detection using individualised baselines and histogram-based descriptors. In2015 IEEE international conference on systems, man, and cybernetics. IEEE, 1864– 1869
2015
-
[8]
Xinqi Fan, Xueli Chen, Mingjie Jiang, Ali Raza Shahid, and Hong Yan. 2023. SelfME: Self-supervised motion learning for micro-expression recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 13834–13843
2023
-
[9]
Xinqi Fan, Jingting Li, John See, Moi Hoon Yap, Wen-Huang Cheng, Xiaobai Li, Xiaopeng Hong, Su-Jing Wang, and Adrian K Davision. 2025. MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering. arXiv preprint arXiv:2506.15298 (2025)
2025
-
[10]
Gunnar Farnebäck. 2003. Two-Frame Motion Estimation Based on Polynomial Expansion. In Image Analysis, Josef Bigun and Tomas Gustavsson (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 363–370
2003
-
[11]
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision. 6202–6211
2019
-
[12]
Yee Siang Gan, Sze-Teng Liong, Wei-Chuen Yau, Yen-Chang Huang, and Lit- Ken Tan. 2019. OFF-ApexNet on micro-expression recognition system. Signal Processing: Image Communication 74 (2019), 129–139
2019
-
[13]
Albert Gu and Tri Dao. 2023. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv preprint arXiv:2312.00752 (2023)
2023 arXiv
-
[14]
Cunhan Guo and Heyan Huang. 2023. Gleffn: A global-local event feature fusion network for micro-expression recognition. In Proceedings of the 3rd Workshop on Facial Micro-Expression: Advanced Techniques for Multi-Modal Facial Expression Analysis. 17–24
2023
-
[15]
Yifei Guo, Bing Li, Xianye Ben, Yi Ren, Junping Zhang, Rui Yan, and Yujun Li
-
[17]
SL Happy and Aurobinda Routray. 2017. Fuzzy histogram of optical flow orienta- tions for micro-expression recognition.IEEE transactions on affective computing 10, 3 (2017), 394–406
2017
-
[20]
in the wild
Petr Husák, Jan Cech, and Jiří Matas. 2017. Spotting facial micro-expressions “in the wild”. In 22nd Computer Vision Winter Workshop (Retz). 1–9
2017
-
[21]
Wenhao Leng, Sirui Zhao, Yiming Zhang, Shiifeng Liu, Xinglong Mao, Hao Wang, Tong Xu, and Enhong Chen. 2022. Abpn: Apex and boundary perception network for micro-and macro-expression spotting. In Proceedings of the 30th ACM International Conference on Multimedia. 7160–7164
2022
-
[22]
Jingting Li, Zizhao Dong, Shaoyuan Lu, Su-Jing Wang, Wen-Jing Yan, Yinhuan Ma, Ye Liu, Changbing Huang, and Xiaolan Fu. 2023. CAS(ME)3: A Third Generation Facial Spontaneous Micro-Expression Database With Depth Information and High Ecological Validity. IEEE Transactions on Pat...
2023
-
[23]
Jingting Li, Catherine Soladie, and Renaud Seguier. 2020. Local temporal pattern and data augmentation for spotting micro-expressions. IEEE Transactions on Affective Computing 14, 1 (2020), 811–822
2020
-
[24]
Jingting Li, Catherine Soladie, Renaud Séguier, Su-Jing Wang, and Moi Hoon Yap
-
[25]
Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. 2024. Videomamba: State space model for efficient video understanding. arXiv preprint arXiv:2403.06977 (2024)
2024 arXiv
-
[26]
Xiaobai Li, Xiaopeng Hong, Antti Moilanen, Xiaohua Huang, Tomas Pfister, Guoying Zhao, and Matti Pietikäinen. 2017. Towards reading hidden emotions: A comparative study of spontaneous micro-expression spotting and recognition methods. IEEE transactions on affective computing 9...
2017
-
[27]
Xiaodong Li, Jiajun Li, Wenchao Du, Hu Chen, and Hongyu Yang. 2024. Learn- ing interval-aware embedding for macro-and micro-expression spotting. In Proceedings of the Asian Conference on Computer Vision. 337–353
2024
-
[28]
Yante Li, Jinsheng Wei, Yang Liu, Janne Kauttonen, and Guoying Zhao. 2022. Deep Learning for Micro-Expression Recognition: A Survey. IEEE Transactions on Affective Computing 13, 4 (2022), 2028–2046. doi:10.1109/TAFFC.2022.3205170
2022
-
[29]
Gen-Bing Liong, Sze-Teng Liong, Chee Seng Chan, and John See. 2024. SFAMNet: A scene flow attention-based micro-expression network. Neurocomputing 566 (2024), 126998
2024
-
[30]
Gen Bing Liong, Sze-Teng Liong, John See, and Chee-Seng Chan. 2022. Mtsn: A multi-temporal stream network for spotting facial macro-and micro-expression with hard and soft pseudo-labels. In Proceedings of the 2nd Workshop on Facial Micro-Expression: Advanced Techniques for Mul...
2022
-
[31]
Gen-Bing Liong, John See, and Chee-Seng Chan. 2023. Spot-then-recognize: A micro-expression analysis network for seamless evaluation of long videos.Signal Processing: Image Communication 110 (2023), 116875
2023
-
[32]
Gen-Bing Liong, John See, and Lai-Kuan Wong. 2021. Shallow Optical Flow Three- Stream CNN For Macro- And Micro-Expression Spotting From Long Videos. In 2021 IEEE International Conference on Image Processing (ICIP). 2643–2647. doi:10.1109/ICIP42928.2021.9506349
2021
-
[33]
Sze-Teng Liong, Y. S. Gan, John See, Huai-Qian Khor, and Yen-Chang Huang. 2019. Shallow Triple Stream Three-dimensional CNN (STSTNet) for Micro-expression Recognition. In 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019). 1–5. doi:10.110...
2019
-
[34]
Sze-Teng Liong, John See, KokSheik Wong, and Raphael Chung-Wei Phan. 2017. Automatic Micro-expression Recognition from Long Video Using a Single Spotted Apex. In Computer Vision – ACCV 2016 Workshops, Chu-Song Chen, Jiwen Lu, and Kai-Kuang Ma (Eds.). Springer International Pub...
2017
-
[35]
Sze-Teng Liong, John See, KokSheik Wong, and Raphael C-W Phan. 2018. Less is more: Micro-expression recognition from video using apex frame. Signal Processing: Image Communication 62 (2018), 82–92
2018
-
[36]
Yong-Jin Liu, Jin-Kai Zhang, Wen-Jing Yan, Su-Jing Wang, Guoying Zhao, and Xiaolan Fu. 2015. A main directional mean optical flow feature for spontaneous micro-expression recognition. IEEE Transactions on Affective Computing 7, 4 (2015), 299–310
2015
-
[37]
Antti Moilanen, Guoying Zhao, and Matti Pietikäinen. 2014. Spotting rapid facial movements from videos using appearance-based feature difference analysis. In 2014 22nd international conference on pattern recognition. IEEE, 1722–1727
2014
-
[38]
Xuan-Bac Nguyen, Chi Nhan Duong, Xin Li, Susan Gauch, Han-Seok Seo, and Khoa Luu. 2023. Micron-bert: Bert-based facial micro-expression recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1482–1492
2023
-
[39]
Timo Ojala, Matti Pietikainen, and David Harwood. 1994. Performance evaluation of texture measures with classification based on Kullback discrimination of distri- butions. In Proceedings of 12th international conference on pattern recognition, Vol. 1. IEEE, 582–585
1994
-
[40]
Hang Pan, Lun Xie, and Zhiliang Wang. 2020. Local bilinear convolutional neural network for spotting macro-and micro-expression intervals in long video sequences. In 2020 15th IEEE international conference on automatic face and gesture recognition (FG 2020). IEEE, 749–753
2020
-
[41]
Devangini Patel, Guoying Zhao, and Matti Pietikäinen. 2015. Spatiotemporal in- tegration of optical flow vectors for micro-expression detection. In International conference on advanced concepts for intelligent vision systems. Springer, 369– 380
2015
-
[42]
John See, Jingting Li, Adrian K Davison, Gen Bing Liong, Moi Hoon Yap, Wen- Huang Cheng, Xiaobai Li, Xiaopeng Hong, and Su-Jing Wang. 2024. MEGC2024: ACM Multimedia 2024 Facial Micro-Expression Grand Challenge. InProceedings of the 32nd ACM International Conference on Multimed...
2024
-
[43]
John See, Moi Hoon Yap, Jingting Li, Xiaopeng Hong, and Su-Jing Wang. 2019. MEGC 2019 – The Second Facial Micro-Expressions Grand Challenge. In 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition MM ’25, October 27–31, 2025, Dublin, Ireland Zizheng ...
2019
-
[44]
Zhiwen Shao, Yifan Cheng, Feiran Li, Yong Zhou, Xuequan Lu, Yuan Xie, and Lizhuang Ma. 2025. MOL: Joint Estimation of Micro-Expression, Optical Flow, and Landmark via Transformer-Graph-Style Convolution. arXiv preprint arXiv:2506.14511 (2025)
2025 arXiv
-
[45]
Matthew Shreve, Jesse Brizzi, Sergiy Fefilatyev, Timur Luguev, Dmitry Goldgof, and Sudeep Sarkar. 2014. Automatic expression spotting in videos. Image and Vision Computing 32, 8 (2014), 476–486
2014
-
[46]
Bo Sun, Siming Cao, Jun He, and Lejun Yu. 2019. Two-stream attention-aware network for spontaneous micro-expression movement spotting. In2019 IEEE 10th International Conference on Software Engineering and Service Science (ICSESS). IEEE, 702–705
2019
-
[47]
Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Raphaël C-W Phan, and Huey-Fang Ong. 2025. Causal-Ex: Causal Graph-based Micro and Macro Expression Spotting. arXiv preprint arXiv:2503.09098 (2025)
2025
-
[48]
Thuong-Khanh Tran, Quang-Nhat Vo, Xiaopeng Hong, and Guoying Zhao
-
[49]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)
2017
-
[50]
Michiel Verburg and Vlado Menkovski. 2019. Micro-expression detection in long videos using optical flow and recurrent neural networks. In 2019 14th IEEE International conference on automatic face & gesture recognition (FG 2019). IEEE, 1–6
2019
-
[51]
Su-Jing Wang, Ying He, Jingting Li, and Xiaolan Fu. 2021. MESNet: A Convo- lutional Neural Network for Spotting Multi-Scale Micro-Expression Intervals in Long Videos. IEEE Transactions on Image Processing 30 (2021), 3956–3969. doi:10.1109/TIP.2021.3064258
2021
-
[52]
In IS&T International Symposium on Electronic Imaging: Science and technology, Imaging and Multimedia Analytics in a Weband Mobile World2019
Dense prediction for micro-expression spotting based on deep sequence model. In IS&T International Symposium on Electronic Imaging: Science and technology, Imaging and Multimedia Analytics in a Weband Mobile World2019. 13-17 Janyuary 2019, Burlingame, USA. Society for Imaging ...
2019
-
[53]
Su-Jing Wang, Wen-Jing Yan, Xiaobai Li, Guoying Zhao, Chun-Guang Zhou, Xiaolan Fu, Minghao Yang, and Jianhua Tao. 2015. Micro-expression recognition using color spaces. IEEE Transactions on Image Processing 24, 12 (2015), 6034– 6047
2015
-
[54]
Mengting Wei, Xingxun Jiang, Wenming Zheng, Yuan Zong, Cheng Lu, and Jiateng Liu. 2023. Cmnet: contrastive magnification network for micro-expression recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 119–127
2023
-
[55]
Zhaoqiang Xia, Xiaoyi Feng, Xiaopeng Hong, and Guoying Zhao. 2018. Spon- taneous facial micro-expression recognition via deep convolutional network. In 2018 eighth international conference on image processing theory, tools and applications (IPTA). IEEE, 1–6
2018
-
[56]
Su-Jing Wang, Shuhang Wu, Xingsheng Qian, Jingxiu Li, and Xiaolan Fu. 2017. A main directional maximal difference analysis for spotting facial movements from long-term videos. Neurocomputing 230 (2017), 382–389
2017
-
[57]
Zhaoqiang Xia, Wei Peng, Huai-Qian Khor, Xiaoyi Feng, and Guoying Zhao. 2020. Revealing the Invisible With Model and Data Shrinking for Composite-Database Micro-Expression Recognition. IEEE Transactions on Image Processing 29 (2020), 8590–8605. doi:10.1109/TIP.2020.3018222
2020
-
[58]
Feng Xu, Junping Zhang, and James Z Wang. 2017. Microexpression identification and categorization using a facial dynamics map. IEEE transactions on affective computing 8, 2 (2017), 254–267
2017
-
[59]
Henian Yang, Shucheng Huang, and Mingxing Li. 2025. MSOF: A main and sec- ondary bi-directional optical flow feature method for spotting micro-expression. Neurocomputing 630 (2025), 129676
2025
-
[60]
Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng, Xianlin Peng, and Guoying Zhao. 2016. Spontaneous micro-expression spotting via geometric deformation modeling. Computer Vision and Image Understanding 147 (2016), 87–94
2016
-
[62]
Chuin Hong Yap, Moi Hoon Yap, Adrian Davison, Connah Kendrick, Jingting Li, Su-Jing Wang, and Ryan Cunningham. 2022. 3D-CNN for Facial Micro- and Macro-expression Spotting on Long Video Sequences using Temporal Oriented Reference Frame. In Proceedings of the 30th ACM Internati...
2022
-
[63]
Jun Yu, Gongpeng Zhao, Yaohui Zhang, Peng He, Zerui Zhang, Zhao Yang, Qing- song Liu, Jianqing Sun, and Jiaen Liang. 2024. Temporal-Informative Adapters in VideoMAE V2 and Multi-Scale Feature Fusion for Micro-Expression Spotting- then-Recognize. In Proceedings of the 32nd ACM ...
2024
-
[64]
Henian Yang, Shucheng Huang, and Mingxing Li. 2025. Ofct: a micro-expression spotting method fusing optical flow features and category text information. Complex & Intelligent Systems 11, 8 (2025), 1–16
2025
-
[65]
Wang-Wang Yu, Jingwen Jiang, Kai-Fu Yang, Hong-Mei Yan, and Yong-Jie Li
-
[66]
Wang-Wang Yu, Kai-Fu Yang, Hong-Mei Yan, and Yong-Jie Li. 2025. Weakly Super- vised Micro-and Macro-Expression Spotting Based on Multi-Level Consistency. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)
2025
-
[67]
He Yuhong. 2021. Research on Micro-Expression Spotting Method Based on Optical Flow Features. InProceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21). Association for Computing Ma- chinery, New York, NY, USA, 4803–4807. doi:10.1145/...
2021
-
[68]
Wang-Wang Yu, Jingwen Jiang, and Yong-Jie Li. 2021. LSSNet: A two-stream convolutional neural network for spotting macro-and micro-expression in long videos. In Proceedings of the 29th ACM International Conference on Multimedia. 4745–4749
2021
-
[69]
Bohao Zhang, Xuejiao Wang, Changbo Wang, and Gaoqi He. 2025. Dynamic Stereotype Theory Induced Micro-expression Recognition with Oriented De- formation. In Proceedings of the Computer Vision and Pattern Recognition Conference. 10701–10711
2025
-
[70]
Li-Wei Zhang, Jingting Li, Su-Jing Wang, Xian-Hua Duan, Wen-Jing Yan, Hai- Yong Xie, and Shu-Cheng Huang. 2020. Spatio-temporal fusion for Macro- and Micro-expression Spotting in Long Video Sequences. In 2020 15th IEEE International Conference on Automatic Face and Gesture Rec...
2020
-
[71]
Shiyu Zhang, Bailan Feng, Zhineng Chen, and Xiangsheng Huang. 2016. Micro-expression recognition by aggregating local spatio-temporal patterns. In International conference on multimedia modeling. Springer, 638–648
2016
-
[72]
Zhihao Zhang, Tong Chen, Hongying Meng, Guangyuan Liu, and Xiaolan Fu
-
[73]
Zhijun Zhai, Jianhui Zhao, Chengjiang Long, Wenju Xu, Shuangjiang He, and Hui- juan Zhao. 2023. Feature representation learning with adaptive displacement gen- eration and transformer fusion for micro-expression recognition. In Proceedings of the IEEE/CVF conference on compute...
2023
-
[74]
Sirui Zhao, Huaying Tang, Xinglong Mao, Shifeng Liu, Yiming Zhang, Hao Wang, Tong Xu, and Enhong Chen. 2023. DFME: A New Benchmark for Dynamic Facial Micro-expression Recognition.IEEE Transactions on Affective Computing (2023), 1–16. doi:10.1109/TAFFC.2023.3341918
2023
-
[75]
Jianxiong Zhou and Ying Wu. 2024. Micro-expression spotting with a novel wavelet convolution magnification network in long videos. Pattern Recognition Letters 178 (2024), 130–137
2024
-
[76]
Ling Zhou, Qirong Mao, Xiaohua Huang, Feifei Zhang, and Zhihong Zhang
-
[77]
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv preprint arXiv:2401.09417 (2024)
2024 arXiv
-
[78]
IEEE Access 6 (2018), 71143–71151
SMEConvNet: A convolutional neural network for spotting spontaneous facial micro-expression from long videos. IEEE Access 6 (2018), 71143–71151
2018
-
[85]
Bochao Zou, Zizheng Guo, Xiaocheng Hu, and Huimin Ma. 2025. Rhythm- Mamba: Fast, Lightweight, and Accurate Remote Physiological Measurement. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 11077– 11085
2025
-
[86]
Bochao Zou, Zizheng Guo, Wenfeng Qin, Xin Li, Kangsheng Wang, and Huimin Ma. 2025. Synergistic Spotting and Recognition of Micro-Expression via Tem- poral State Transition. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)....
2025
-
[2018]
doi:10.1109/TAFFC.2016.2573832
SAMM: A Spontaneous Micro-Facial Movement Dataset.IEEE Transactions on Affective Computing 9, 1 (2018), 116–129. doi:10.1109/TAFFC.2016.2573832
2018
-
[2019]
In 2019 14th IEEE International conference on automatic face & gesture recognition (FG 2019)
Spotting micro-expressions on long videos sequences. In 2019 14th IEEE International conference on automatic face & gesture recognition (FG 2019). IEEE, 1–5
2019
-
[2021]
IEEE MultiMedia 28, 2 (2021), 29–39
A magnitude and angle combined optical flow feature for microexpression spotting. IEEE MultiMedia 28, 2 (2021), 29–39
2021
-
[2022]
doi:10.1016/j.patcog.2021.108275
Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition.Pattern Recognition 122 (2022), 108275. doi:10.1016/j.patcog.2021.108275
2022
-
[2024]
IEEE Transactions on Affective Computing 15, 1 (2024), 223–240
LGSNet: A Two-Stream Network for Micro- and Macro-Expression Spotting With Background Modeling. IEEE Transactions on Affective Computing 15, 1 (2024), 223–240. doi:10.1109/TAFFC.2023.3266808
2024
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.