Pith. sign in

REVIEW 5 cited by

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.16566 v2 pith:K3KYQ6B2 submitted 2025-01-27 cs.HC

classification cs.HC
keywords emotionunderstandingaffectgptdatasetlanguagemultimodalbenchmarkmllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level, from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

    cs.HC 2026-06 conditional novelty 7.0 of 10

    COSI-Lab is a weakly scripted conference-workshop dataset with multi-perspective apparent-intent annotations, self-reported goals, and benchmarks for social intention inference and conversation group detection.

  2. DeceptionX: From Multimodal Evidence to Explainable Deception Detection

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    DeceptionX trains a multimodal LLM to explain lie judgments with visual and audio evidence, but its benchmark gains rest on overlapping train/test data and label-conditioned annotations.

  3. Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new large-scale facial emotion caption dataset and a global-local contrastive training framework with positive mining improve zero-shot facial expression recognition.

  4. Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony

    cs.HC 2026-07 conditional novelty 5.0 of 10

    Group-level EEG dynamic neural synchrony (CorrCA) preferentially tracks the rate of change of continuous arousal and shows valence-dependent structure across four datasets.

  5. EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

    cs.AI 2026-02 conditional novelty 5.0 of 10

    EMO-R3, which combines a three-step emotional reasoning prompt with a reward for the model agreeing with its own image–emotion judgments, raises visual emotion-recognition accuracy by about one point over plain GRPO.

Pith tools