REVIEW 4 major objections 4 minor 33 references
The Importance of Facial Features in Vision-based Sign Language Recognition: Eyes, Mouth or Full Face?
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Mouth, not eyes, is the facial feature sign recognition needs.
desk verdict The mouth-over-eyes ranking is plausible, but the 'body' baseline likely contains the face, so the central claim needs a cleaner control before it can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The experimental design is the mechanism: the same video clip is cropped into four regions of interest (eyes, mouth, full face, and upper body including hands), and each region is fed, alone or in a late-fusion two-stream configuration, into either a channel-separated convolutional network or a multiscale vision transformer. Because the only change between runs is which pixels are visible, any accuracy difference is attributed to that facial region. Gradient-based saliency maps (vanilla gradients plus SmoothGrad) provide the qualitative counterpart, showing where the trained models actually look.
What would settle it
Retrain the body-only and body+mouth models with a signer-disjoint split and with the face explicitly masked out of the body crop; if body+mouth no longer beats body alone, or the gap largely disappears, the mouth's central role is an artifact of leaked facial pixels or signer appearance.
Extended reading notes
Core claim
On its own terms, the paper establishes that in an end-to-end, RGB-video, gloss-level recognition setting, the mouth is the most informative facial region, and that the value of the full face comes largely from the mouth area. With 12 randomly selected glosses from a public German Sign Language corpus, a channel-separated convolutional network (CSN) and a multiscale vision transformer (MViT) both show the same pattern: mouth-only inputs outperform eye-only inputs by a wide margin; fusing mouth with the body stream improves accuracy over body alone (CSN: 80.53% to 88.24% top-1) while fusing eyes does not; and body+mouth and body+face are statistically indistinguishable across all metrics. The authors interpret this as evidence that mouth actions, which can distinguish signs with identical manual articulation, are the primary non-manual facial contribution, and they confirm it with gradient saliency maps that consistently highlight the mouth.
Load-bearing premise
The comparison assumes the body crop is a manual-only baseline, yet the cropping procedure never explicitly removes facial pixels from it, so the “manual” stream may already contain mouth and eye information; the random clip split also assumes no signer appears in both training and test.
Editorial extensions
If this is right
- For isolated sign recognition, systems that already use manual features should include a mouth stream; an eye stream is unlikely to add accuracy.
- Fusing body and mouth yields essentially the same accuracy as fusing body and full face, so a small mouth crop appears to retain most of the facial benefit at lower computational cost.
- Facial information mainly helps the model decide between its top candidates, since top-3 accuracy shows the largest gains, consistent with mouthing disambiguating manually similar signs.
- Because the pattern holds in both a convolutional and a transformer-based architecture, the result is not specific to one model family.
Reading between the lines
- Editorial: the “body” ROI was cropped from upper-body and hand landmarks without explicitly removing the face, so facial pixels may leak into the manual-only baseline; a cleaner manual baseline would mask or crop out the face before testing how much the mouth adds.
- Editorial: the 8:1:1 clip split is described only as random, not as signer-disjoint; if clips from the same signer appear in both training and test, signer appearance rather than mouthing could contribute. A signer-independent split would test this.
- Editorial: a direct ablation that masks or blurs the mouth in full-face inputs, or swaps mouth regions between signers, would test whether mouthing itself, rather than head motion or identity, drives the gain.
- Editorial: pretraining on lip-reading before fine-tuning on sign language is a natural next test, since the paper points to lip reading as a possible source of transferable features.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates the contribution of facial regions (eyes, mouth, full face) to vision-based isolated sign language recognition. The authors construct a 12-class German Sign Language dataset from the Public DGS Corpus, crop videos into four regions of interest (eyes, mouth, face, body), train a CNN-based model (CSN) and a transformer-based model (MViT) on each ROI, and also evaluate late fusion of the body ROI with each facial ROI. They report top-1/top-3 accuracy and F1 scores, supplemented by SmoothGrad saliency maps. The central claim is that the mouth is the most important non-manual facial feature and that incorporating facial features, especially the mouth, improves recognition.
Significance. If the claims are valid, the paper offers a useful controlled comparison of facial-feature contributions across two architecturally distinct models, which is rare in the sign-language-recognition literature. Strengths include the use of a public corpus, a consistent late-fusion protocol applied identically to both models, the explicit framing of the study as a controlled analysis rather than a state-of-the-art contribution, and the inclusion of qualitative saliency evidence. The conclusion that mouth cues are the most informative facial region is plausible and consistent with linguistic and eye-tracking studies. However, the central quantitative claim depends on a body baseline whose construction is not shown to exclude facial pixels, and the paper's statistical and dataset-split choices weaken the stated conclusions.
major comments (4)
- [Section 3 (Dataset), body ROI definition] The body ROI is defined as the bounding box around MediaPipe upper-body and hand landmarks, but the paper does not state that facial landmarks are excluded from this bounding box. MediaPipe's upper-body pose landmark set includes facial landmarks such as the nose, eyes, ears, and mouth corners, so an axis-aligned box around these landmarks will generally contain the face. If the body condition already contains facial pixels, then the central comparison between 'body' and 'body + facial ROI' is not a comparison of manual features versus manual plus facial features: the observed gains could reflect the addition of a second, higher-resolution facial crop rather than the addition of new facial information. Moreover, the equivalence between body+mouth and body+face, which is used to conclude that the mouth is the most important facial feature, could be an artifact of the full face being largely redundant with facial pixels already present in the body crop. The authors should recompute the body ROI after explicitly excluding facial landmarks, report overlap statistics between the body and face crops, or add a control condition with the face masked or pixelated in the body stream. Without such a control, the main claim is not cleanly supported.
- [Section 3 (Dataset), split paragraph] The 8:1:1 split is described only as random while preserving class distribution; it is not stated whether clips from the same signer are kept in the same set. Since the DGS Corpus contains multiple clips per signer and the facial ROIs carry strong identity cues, a random clip split can allow the model to exploit signer appearance rather than sign identity. This could inflate the apparent contribution of face and mouth streams and affect the feature ranking. The authors should specify whether the split is signer-independent, report the degree of signer overlap between train and test, or re-run the experiments with a signer-exclusive split. If such a split is not feasible, the conclusions should be explicitly restricted to within-signer recognition rather than general ASLR.
- [Section 5.1, Table 1] The paper uses overlapping 95% confidence intervals to conclude that body+eyes is 'not significantly different' from body alone and that body+mouth and body+face are 'statistically indistinguishable'. Overlap of confidence intervals is not an equivalence test, and a non-significant difference is not evidence of absence of difference. The paper also does not report how the confidence intervals were computed (number of training runs or seeds, bootstrap over clips). Because the mouth-versus-face equivalence is load-bearing for the conclusion that the mouth is the most important facial feature, the authors should either provide a paired statistical test or equivalence bounds, or weaken the conclusion to 'we found no significant difference in this setting' and explicitly treat it as a null result.
- [Abstract and Section 5.1] The abstract and conclusion state that the mouth 'significantly improving accuracy' and that incorporating facial features significantly improves recognition, but Table 1 shows that for MViT the top-1 gains of body+mouth (86.42±2.51) and body+face (86.98±2.47) over body (84.03±2.69) are not significant by the paper's own confidence-interval criterion, and Section 5.1 explicitly says so. The unqualified wording overstates the evidence. The claims should be qualified by model and metric, e.g., significant for CSN top-1 accuracy and MViT top-3 accuracy.
minor comments (4)
- [Section 5.1] The sentence 'the mouth area is the most important facial feature .' contains a stray space before the period; please correct the typographical error.
- [Section 3 (Dataset)] The paper says the 12 glosses were 'randomly selected' but does not report the random seed or the full list of selected glosses; providing these details would improve reproducibility.
- [Section 4 (Experiments)] The experimental setup does not state the number of random seeds or training repetitions used to compute the confidence intervals in Table 1; please specify this for reproducibility.
- [Section 5.2 (Saliency Maps)] The saliency analysis is described qualitatively; please clarify the normalization procedure for the attribution values and state how many videos/classes were inspected and whether any quantitative measure of mouth-region saliency was computed.
Circularity Check
No significant circularity: the central mouth-importance claim rests on the paper's own controlled ablations and confidence-interval comparisons, not on fitted predictions or self-cited results.
full rationale
This paper is an empirical ablation study rather than a derivation from first principles, so most derivation-style circularity patterns do not apply. The central claim that the mouth is the most important facial feature is supported by the authors' own experimental comparisons in Table 1: single-ROI accuracies are used only descriptively, while the mouth-vs-face conclusion follows from overlapping confidence intervals between body+mouth and body+face across two architectures. No fitted parameter is renamed as a prediction, and no equation reduces a result to its own input. The self-citations to [25] are contextual: [25] is cited to note that mouth actions can disambiguate signs with similar manual articulation, which is background linguistic motivation and not the evidence for the paper's ranking. The saliency-map analysis does use the same trained models whose outputs it is said to confirm, but the paper explicitly treats it as a qualitative complement to the quantitative ablation, not as an independent benchmark; internal corroboration of this kind is not a circular reduction. The possible inclusion of facial pixels in the body ROI and the lack of a signer-independent split are experimental validity concerns about the manual-vs-facial contrast, but they do not make the paper's conclusion equivalent to its inputs by construction. No quoted step exhibits the specific type of reduction required to establish circularity, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (3)
- Class frequency threshold and number of glosses =
12 glosses, each with >=500 occurrences in the Public DGS Corpus
- Video length normalization =
32 frames with last-frame padding
- Training hyperparameters =
learning rate 1e-5, batch size 3, 100 epochs, RandAugment N=3 magnitude 4
assumptions (4)
- domain assumption Kinetics-400 pretrained CSN and MViT transfer to isolated German Sign Language recognition
- domain assumption Saliency maps computed with vanilla gradients plus SmoothGrad reflect the features the model actually uses
- domain assumption Overlapping 95% confidence intervals are a valid test of whether two conditions differ
- ad hoc to paper The random 8:1:1 clip split does not leak signer identity between training and test
Cite this review
Pith. "Pith review of The Importance of Facial Features in Vision-based Sign Language Recognition: Eyes, Mouth or Full Face?." pith.science (2026). https://pith.science/paper/JGWCYS7E
@misc{pith2026250720884,
author = {Pith},
title = {Pith review of: The Importance of Facial Features in Vision-based Sign Language Recognition: Eyes, Mouth or Full Face?},
year = {2026},
howpublished = {\url{https://pith.science/paper/JGWCYS7E}},
note = {Machine review of arXiv:2507.20884}
}
read the original abstract
Non-manual facial features play a crucial role in sign language communication, yet their importance in automatic sign language recognition (ASLR) remains underexplored. While prior studies have shown that incorporating facial features can improve recognition, related work often relies on hand-crafted feature extraction and fails to go beyond the comparison of manual features versus the combination of manual and facial features. In this work, we systematically investigate the contribution of distinct facial regionseyes, mouth, and full faceusing two different deep learning models (a CNN-based model and a transformer-based model) trained on an SLR dataset of isolated signs with randomly selected classes. Through quantitative performance and qualitative saliency map evaluation, we reveal that the mouth is the most important non-manual facial feature, significantly improving accuracy. Our findings highlight the necessity of incorporating facial features in ASLR.
Figures
Reference graph
Works this paper leans on
-
[1]
Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh. 2021. OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields. IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 1 (2021), 172–
work page 2021
-
[2]
Ekin Dogus Cubuk, Barret Zoph, Jon Shlens, and Quoc Le. 2020. RandAug- ment: Practical Automated Data Augmentation with a Reduced Search Space. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ran- zato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 18613–18624
work page 2020
-
[3]
Karen Emmorey, Robin Thompson, and Rachael Colvin. 2009. Eye gaze during comprehension of American Sign Language by native and beginning signers. J. Deaf Stud. Deaf Educ. 14, 2 (2009), 237–243
work page 2009
-
[4]
Haoqi Fan, Tullie Murrell, Heng Wang, Kalyan Vasudev Alwala, Yanghao Li, Yilei Li, Bo Xiong, Nikhila Ravi, Meng Li, Haichuan Yang, Jitendra Malik, Ross Girshick, Matt Feiszli, Aaron Adcock, Wan-Yen Lo, and Christoph Feichtenhofer
-
[5]
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. 2021. Multiscale Vision Transformers. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . 6804–6815. doi:10.1109/ICCV48922.2021.00675
arXiv 2021
-
[6]
Jérôme Fink, Benoît Frénay, Laurence Meurant, and Anthony Cleve. 2021. LSFB- CONT and LSFB-ISOL: Two New Datasets for Vision-Based Sign Language Recognition. In 2021 International Joint Conference on Neural Networks (IJCNN) . 1–8. doi:10.1109/IJCNN52387.2021.9534336
arXiv 2021
-
[7]
Ye Gao, Ruixiang Hu, Tian Ma, Songyi Guo, Yizhou Yang, and Xinlei Zhou. 2022. Dynamic Sign Language Recognition Based on Improved R(2+1)D Algorithm. In 2022 7th International Conference on Image, Vision and Computing (ICIVC) . 7–15. doi:10.1109/ICIVC55077.2022.9886615
arXiv 2022
-
[8]
Xiangzu Han, Fei Lu, Jianqin Yin, Guohui Tian, and Jun Liu. 2022. Sign Language Recognition Based on R(2+1)D With Spatial–Temporal–Channel Attention. IEEE Transactions on Human-Machine Systems 52, 4 (2022), 687–698. doi:10.1109/THMS. 2022.3144000
arXiv 2022
Show all 33 references
-
[9]
Thomas Hanke, Marc Schulder, Reiner Konrad, and Elena Jahn. 2020. Extending the Public DGS Corpus in Size and Depth. In Proceedings of the LREC2020 9th Workshop on the Representation and Processing of Sign Languages: Sign Language Resources in the Service of the Language Commu...
2020
-
[10]
Annika Herrmann. 2014. Modal and Focus Particles in Sign Languages: A Cross-Linguistic Study. De Gruyter Mouton, Berlin, Boston. doi:doi:10.1515/ 9781614511816
2014
-
[11]
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman. 2017. The Kinetics Human Action Video Dataset. arXiv preprint arXiv:1705.06950 (2017)
2017 arXiv
-
[12]
Oscar Koller. 2020. Quantitative Survey of the State of the Art in Sign Language Recognition. arXiv preprint arXiv:2008.09918 (2020)
2020 arXiv
-
[13]
Reiner Konrad, Thomas Hanke, Gabriele Langer, Dolly Blanck, Julian Bleicken, Ilona Hofmann, Olga Jeziorski, Lutz König, Susanne König, Rie Nishio, Anja Regen, Uta Salden, Sven Wagner, Satu Worseck, Oliver Böse, Elena Jahn, and Marc Schulder. 2020. MEINE DGS – annotiert. Öffent...
2020 doi
-
[14]
Reiner Konrad, Thomas Hanke, Gabriele Langer, Susanne König, Lutz König, Rie Nishio, and Anja Regen. 2022. Öffentliches DGS-Korpus: Annotationskonventionen / Public DGS Corpus: Annotation Conventions (4.1 ed.). Project Note AP03-2018-01. DGS-Korpus project, IDGS, Hamburg Unive...
2022
-
[15]
Bartosz Krawczyk. 2016. Learning from imbalanced data: open challenges and future directions. Progress in Artificial Intelligence 5, 4 (April 2016), 221–232. doi:10.1007/s13748-016-0094-0
2016 doi
-
[16]
Pradeep Kumar, Partha Pratim Roy, and Debi Prosad Dogra. 2018. Independent Bayesian classifier combination based sign language recognition using facial expression. Information Sciences 428 (2018), 30–48. doi:10.1016/j.ins.2017.10.046
2018 doi
-
[17]
Dongxu Li, Cristian Rodriguez-Opazo, Xin Yu, and Hongdong Li. 2019. Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison. 2020 IEEE Winter Conference on Applications of Computer Vision (W ACV)(2019), 1448–1458
2019
-
[18]
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, Wan-Teh Chang, Wei Hua, Manfred Georg, and Matthias Grundmann. 2019. MediaPipe: A Framework for Building Perception Pipelines. ar...
2019 arXiv
-
[19]
Pingchuan Ma, Stavros Petridis, and Maja Pantic. 2022. Visual speech recognition for multiple languages in the wild. Nature Machine Intelligence 4, 11 (Oct. 2022), 930–939. doi:10.1038/s42256-022-00550-z
2022 doi
-
[20]
Rodríguez-Ortiz
Eliana Mastrantuono, David Saldaña, and Isabel R. Rodríguez-Ortiz. 2017. An Eye Tracking Study on the Perception and Comprehension of Unimodal and Bimodal Linguistic Inputs by Deaf Adolescents. Frontiers in Psychology Volume 8 - 2017 (2017). doi:10.3389/fpsyg.2017.01044
2017
-
[21]
Medet Mukushev, Arman Sabyrov, Alfarabi Imashev, Kenessary Koishybay, Vadim Kimmelman, and Anara Sandygulova. 2020. Evaluation of Manual and Non- manual Components for Sign Language Recognition. In 12th International Con- ference on Language Resources and Evaluation (LREC 2020...
2020
-
[22]
Nina-Kristin Pendzich. 2020. Lexical Nonmanuals in German Sign Language: Empirical Studies and Theoretical Implications. De Gruyter Mouton, Berlin, Boston. doi:doi:10.1515/9783110671667
2020 doi
-
[23]
López-Ortiz, Luis M
Marina Perea-Trigo, Enrique J. López-Ortiz, Luis M. Soria-Morillo, Juan A. Ál- varez García, and J. J. Vegas-Olmos. 2024. Impact of face swapping and data augmentation on sign language recognition. Universal Access in the Information Society (July 2024). doi:10.1007/s10209-024-01133-y
2024 doi
-
[24]
Dinh Nam Pham and Eleftherios Avramidis. 2025. Transfer Learning from Visual Speech Recognition to Mouthing Recognition in German Sign Language. In 2025 19th IEEE International Conference on Automatic Face and Gesture Recognition (FG). 1–6
2025
-
[25]
Dinh Nam Pham, Vera Czehmann, and Eleftherios Avramidis. 2023. Disam- biguating Signs: Deep Learning-based Gloss-level Classification for German Sign Language by Utilizing Mouth Actions. In 31st European Symposium on Artificial Neural Networks, Computational Intelligence and M...
2023 doi
-
[26]
Noha Sarhan and Simone Frintrop. 2020. Transfer Learning For Videos: From Action Recognition To Sign Language Recognition. In 2020 IEEE International Conference on Image Processing (ICIP) . 1811–1815. doi:10.1109/ICIP40778.2020. 9191289
2020
-
[27]
Noha Sarhan and Simone Frintrop. 2023. Unraveling a Decade: A Comprehensive Survey on Isolated Sign Language Recognition. In 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) . 3202–3211. doi:10.1109/ ICCVW60793.2023.00345
2023
-
[28]
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Workshop Tra...
2014
-
[29]
Viégas, and Martin Wat- tenberg
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda B. Viégas, and Martin Wat- tenberg. 2017. SmoothGrad: removing noise by adding noise. arXiv preprint arXiv:1706.03825 (2017)
2017 arXiv
-
[30]
Du Tran, Heng Wang, Matt Feiszli, and Lorenzo Torresani. 2019. Video Clas- sification With Channel-Separated Convolutional Networks. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV). 5551–5560. doi:10.1109/ICCV. 2019.00565
2019
-
[31]
Ulrich von Agris, Moritz Knorr, and Karl-Friedrich Kraiss. 2008. The significance of facial features for automatic sign language recognition. In 2008 8th IEEE International Conference on Automatic Face & Gesture Recognition . 1–6. doi:10. 1109/AFGR.2008.4813472 Preprint — Auth...
2008
-
[186]
doi:10.1109/TPAMI.2019.2929257
2019
-
[2021]
In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21)
PyTorchVideo: A Deep Learning Library for Video Understanding. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21). Association for Computing Machinery, New York, NY, USA, 3783–3786. doi:10.1145/3474085.3478329
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.