Pith. sign in

REVIEW 10 cited by

Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.04855 v1 pith:BWFS4AKR submitted 2019-09-25 cs.CV cs.HCcs.LGeess.IV

classification cs.CVcs.HCcs.LGeess.IV
keywords aff-wild2recognitionavailabledatabaseemotionnetworksactiondatabases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Affective computing has been largely limited in terms of available data resources. The need to collect and annotate diverse in-the-wild datasets has become apparent with the rise of deep learning models, as the default approach to address any computer vision task. Some in-the-wild databases have been recently proposed. However: i) their size is small, ii) they are not audiovisual, iii) only a small part is manually annotated, iv) they contain a small number of subjects, or v) they are not annotated for all main behavior tasks (valence-arousal estimation, action unit detection and basic expression classification). To address these, we substantially extend the largest available in-the-wild database (Aff-Wild) to study continuous emotions such as valence and arousal. Furthermore, we annotate parts of the database with basic expressions and action units. As a consequence, for the first time, this allows the joint study of all three types of behavior states. We call this database Aff-Wild2. We conduct extensive experiments with CNN and CNN-RNN architectures that use visual and audio modalities; these networks are trained on Aff-Wild2 and their performance is then evaluated on 10 publicly available emotion databases. We show that the networks achieve state-of-the-art performance for the emotion recognition tasks. Additionally, we adapt the ArcFace loss function in the emotion recognition context and use it for training two new networks on Aff-Wild2 and then re-train them in a variety of diverse expression recognition databases. The networks are shown to improve the existing state-of-the-art. The database, emotion recognition models and source code are available at http://ibug.doc.ic.ac.uk/resources/aff-wild2.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DVD: A Comprehensive Dataset for Advancing Violence Detection in Real-World Scenarios

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The authors propose a new frame-level annotated violence detection dataset, DVD, with 500 videos and rich metadata, but it is not yet available and lacks validation experiments.

  2. Interpretable Concept-based Deep Learning Framework for Multimodal Human Behavior Modeling

    cs.CV 2025-02 conditional novelty 6.0 of 10

    AGCM classifies facial expressions and conversational engagement through learnable, spatially localized human-readable concepts, reporting accuracy at or above black-box baselines on RAF-DB, AffectNet, Aff-Wild2, and NOXI.

  3. AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Adding a conditional rectified-flow head to a DINOv3 multi-task affect model improves valence-arousal CCC by +0.058 when the backbone is frozen and, with fine-tuning and validation-tuned calibration, reaches P_MTL=1.1...

  4. Causal Supervision of Attention for Affective Behaviour Analysis

    cs.CV 2026-07 unverdicted novelty 5.0 of 10

    Causal supervision plus K-V independence and SwiGLU attention pooling yields a multi-task P-score of 1.2214 on s-Aff-Wild2 validation for valence-arousal, expression, and action-unit prediction.

  5. Strength-Parity Ensembling with Parameter-Isolated Experts for Multi-Task Affect Recognition

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Parameter-isolated LoRA experts on one face backbone stay decorrelated (0.91 vs 0.98 for full fine-tuning) and improve an ABAW affect ensemble from 1.6669 to 1.6949–1.7259 validation score.

  6. A Shared Latent for Partially-Labeled Multi-Task Facial Affect Recognition

    cs.CV 2026-07 accept novelty 5.0 of 10

    A shared variational affect latent that marginalizes missing labels lifts rare expression and action-unit recognition on s-Aff-Wild2 beyond masked-loss training.

  7. Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation

    cs.CV 2026-03 reject novelty 5.0 of 10

    Distance-aware soft prompts over a 3×3 emotion grid with CLIP text prototypes and audio-visual GRU fusion achieve CCC_mean 0.5361 on Aff-Wild2, beating only the paper's self-defined baselines.

  8. AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A multi-task affective system combining frozen AffectNet backbones, LoRA-adapted MAE for action units, temporal heads, fusion, and ensembling attains P=1.7302 on the s-Aff-Wild2 validation split.

  9. Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding

    cs.CV 2025-07 reject novelty 4.0 of 10

    A GRU-based cross-attention fusion of frozen vision-language encoders is claimed to achieve strong results on DVD and Aff-Wild2, but the supporting experiments are missing from the paper.

  10. TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation

    cs.MM 2025-07 reject novelty 4.0 of 10

    TAGF adds a BiLSTM-based gate that reweights recursive cross-attention outputs for valence-arousal prediction, with results slightly below several existing methods on Aff-Wild2.

Pith tools