REVIEW 2 cited by
MVT: Mask Vision Transformer for Facial Expression Recognition in the wild
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Facial Expression Recognition (FER) in the wild is an extremely challenging task in computer vision due to variant backgrounds, low-quality facial images, and the subjectiveness of annotators. These uncertainties make it difficult for neural networks to learn robust features on limited-scale datasets. Moreover, the networks can be easily distributed by the above factors and perform incorrect decisions. Recently, vision transformer (ViT) and data-efficient image transformers (DeiT) present their significant performance in traditional classification tasks. The self-attention mechanism makes transformers obtain a global receptive field in the first layer which dramatically enhances the feature extraction capability. In this work, we first propose a novel pure transformer-based mask vision transformer (MVT) for FER in the wild, which consists of two modules: a transformer-based mask generation network (MGN) to generate a mask that can filter out complex backgrounds and occlusion of face images, and a dynamic relabeling module to rectify incorrect labels in FER datasets in the wild. Extensive experimental results demonstrate that our MVT outperforms state-of-the-art methods on RAF-DB with 88.62%, FERPlus with 89.22%, and AffectNet-7 with 64.57%, respectively, and achieves a comparable result on AffectNet-8 with 61.40%.
Forward citations
Cited by 2 Pith papers
-
FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
FEALLM is a multimodal LLM fine-tuned on a new, aligned facial expression and action unit reasoning dataset, reporting improved facial emotion analysis on its benchmark and zero-shot gains on RAF-DB, AffectNet, BP4D, ...
-
Multimodal Prompt Alignment for Facial Expression Recognition
A frozen-CLIP facial expression recognition framework with LLM-guided prompts, prototype regularization, and sparse global-local alignment claims new state-of-the-art accuracy across RAF-DB, FERPlus, and AffectNet.
Discussion (0). Continue with ORCID to comment.