Pith. sign in

REVIEW 2 cited by

MVT: Mask Vision Transformer for Facial Expression Recognition in the wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.04520 v2 pith:AM34SLRV submitted 2021-06-08 cs.CV

classification cs.CV
keywords maskvisionwildfacialtransformerbackgroundsdatasetsexpression
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Facial Expression Recognition (FER) in the wild is an extremely challenging task in computer vision due to variant backgrounds, low-quality facial images, and the subjectiveness of annotators. These uncertainties make it difficult for neural networks to learn robust features on limited-scale datasets. Moreover, the networks can be easily distributed by the above factors and perform incorrect decisions. Recently, vision transformer (ViT) and data-efficient image transformers (DeiT) present their significant performance in traditional classification tasks. The self-attention mechanism makes transformers obtain a global receptive field in the first layer which dramatically enhances the feature extraction capability. In this work, we first propose a novel pure transformer-based mask vision transformer (MVT) for FER in the wild, which consists of two modules: a transformer-based mask generation network (MGN) to generate a mask that can filter out complex backgrounds and occlusion of face images, and a dynamic relabeling module to rectify incorrect labels in FER datasets in the wild. Extensive experimental results demonstrate that our MVT outperforms state-of-the-art methods on RAF-DB with 88.62%, FERPlus with 89.22%, and AffectNet-7 with 64.57%, respectively, and achieves a comparable result on AffectNet-8 with 61.40%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    FEALLM is a multimodal LLM fine-tuned on a new, aligned facial expression and action unit reasoning dataset, reporting improved facial emotion analysis on its benchmark and zero-shot gains on RAF-DB, AffectNet, BP4D, ...

  2. Multimodal Prompt Alignment for Facial Expression Recognition

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A frozen-CLIP facial expression recognition framework with LLM-guided prompts, prototype regularization, and sparse global-local alignment claims new state-of-the-art accuracy across RAF-DB, FERPlus, and AffectNet.

Pith tools