Pith. sign in

REVIEW 2 cited by

From Text to Emotion: Unveiling the Emotion Annotation Capabilities of LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.17026 v1 pith:V7WTH2O2 submitted 2024-08-30 cs.CL

classification cs.CL
keywords emotionhumanannotationannotationsgpt-4llmsmodelstraining
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Training emotion recognition models has relied heavily on human annotated data, which present diversity, quality, and cost challenges. In this paper, we explore the potential of Large Language Models (LLMs), specifically GPT4, in automating or assisting emotion annotation. We compare GPT4 with supervised models and or humans in three aspects: agreement with human annotations, alignment with human perception, and impact on model training. We find that common metrics that use aggregated human annotations as ground truth can underestimate the performance, of GPT-4 and our human evaluation experiment reveals a consistent preference for GPT-4 annotations over humans across multiple datasets and evaluators. Further, we investigate the impact of using GPT-4 as an annotation filtering process to improve model training. Together, our findings highlight the great potential of LLMs in emotion annotation tasks and underscore the need for refined evaluation methodologies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Third-parties Read Our Emotions?

    cs.CL 2025-04 conditional novelty 5.0 of 10

    Third-party emotion annotations, both human and LLM, show low to fair agreement with authors' self-reported emotions, with LLMs outperforming humans but still misaligning substantially.

  2. Affect Models Have Weak Generalizability to Atypical Speech

    cs.LG 2025-04 conditional novelty 5.0 of 10

    Affect models predict sadness more often and neutrality less often for atypical speech, but the reported fine-tuning gain is evaluated against the same pseudo-label source used to train it.

Pith tools