Pith. sign in

REVIEW 3 cited by

ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14952 v3 pith:HVQWMY2T submitted 2024-06-21 cs.CL

classification cs.CL
keywords llmsmodelsrole-playinghumanesc-evalevaluationagentai-assistant
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Emotion Support Conversation (ESC) is a crucial application, which aims to reduce human stress, offer emotional guidance, and ultimately enhance human mental and physical well-being. With the advancement of Large Language Models (LLMs), many researchers have employed LLMs as the ESC models. However, the evaluation of these LLM-based ESCs remains uncertain. Inspired by the awesome development of role-playing agents, we propose an ESC Evaluation framework (ESC-Eval), which uses a role-playing agent to interact with ESC models, followed by a manual evaluation of the interactive dialogues. In detail, we first re-organize 2,801 role-playing cards from seven existing datasets to define the roles of the role-playing agent. Second, we train a specific role-playing model called ESC-Role which behaves more like a confused person than GPT-4. Third, through ESC-Role and organized role cards, we systematically conduct experiments using 14 LLMs as the ESC models, including general AI-assistant LLMs (ChatGPT) and ESC-oriented LLMs (ExTES-Llama). We conduct comprehensive human annotations on interactive multi-turn dialogues of different ESC models. The results show that ESC-oriented LLMs exhibit superior ESC abilities compared to general AI-assistant LLMs, but there is still a gap behind human performance. Moreover, to automate the scoring process for future ESC models, we developed ESC-RANK, which trained on the annotated data, achieving a scoring performance surpassing 35 points of GPT-4. Our data and code are available at https://github.com/AIFlames/Esc-Eval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ESC-Judge automates head-to-head evaluation of emotional-support chatbots using Hill's Exploration-Insight-Action rubric, with an LLM judge reported to match PhD annotators on roughly 85 percent of comparisons.

  2. EmoAssist: Emotional Assistant for Visual Impairment Community

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A fine-tuned LLaVA model, trained with preference optimization on 800 emotional image-QA examples, beats GPT-4o on a new empathy-focused benchmark for visual impairment assistance.

  3. Multimodal Large Language Models for Medicine: A Comprehensive Survey

    cs.LG 2025-04 conditional novelty 2.0 of 10

    A comprehensive review cataloging medical MLLMs, their uses in report generation, diagnosis, and treatment, and the challenges of accuracy, hallucination, fairness, privacy, and deployment.

Pith tools