Pith. sign in

REVIEW 4 cited by

NR-DFERNet: Noise-Robust Network for Dynamic Facial Expression Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.04975 v1 pith:JUUKGTIS submitted 2022-06-10 cs.CV

classification cs.CV
keywords framesdynamicfeaturesexpressionfacialnoisynr-dfernetrecognition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Dynamic facial expression recognition (DFER) in the wild is an extremely challenging task, due to a large number of noisy frames in the video sequences. Previous works focus on extracting more discriminative features, but ignore distinguishing the key frames from the noisy frames. To tackle this problem, we propose a noise-robust dynamic facial expression recognition network (NR-DFERNet), which can effectively reduce the interference of noisy frames on the DFER task. Specifically, at the spatial stage, we devise a dynamic-static fusion module (DSF) that introduces dynamic features to static features for learning more discriminative spatial features. To suppress the impact of target irrelevant frames, we introduce a novel dynamic class token (DCT) for the transformer at the temporal stage. Moreover, we design a snippet-based filter (SF) at the decision stage to reduce the effect of too many neutral frames on non-neutral sequence classification. Extensive experimental results demonstrate that our NR-DFERNet outperforms the state-of-the-art methods on both the DFEW and AFEW benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    MiRA is a parameter-free frame-marginal attention redistribution technique for ViT video models that improves sensitivity to localized facial cues on FER benchmarks.

  2. From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    GRACE pairs motion-weighted video tokens with AI-refined emotion text tokens using optimal transport, reporting new UAR and WAR records on DFEW, FERV39k, and MAFW.

  3. Action Unit Enhance Dynamic Facial Expression Recognition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AU-DFER boosts dynamic facial expression recognition by roughly 1% WAR/UAR through an AU-expression knowledge matrix injected as a weighted AU loss, at no extra inference cost.

  4. Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition

    cs.CV 2025-11 conditional novelty 4.0 of 10

    TG-DFER combines CLIP text prompts, a visual-prompt attention module, and a fine/coarse temporal transformer to nudge DFER benchmarks up slightly: 60.17/71.62 on DFEW and 41.50/51.67 on FERV39k.

Pith tools