REVIEW 6 cited by
Towards a Client-Centered Assessment of LLM Therapists by Client Simulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Although there is a growing belief that LLMs can be used as therapists, exploring LLMs' capabilities and inefficacy, particularly from the client's perspective, is limited. This work focuses on a client-centered assessment of LLM therapists with the involvement of simulated clients, a standard approach in clinical medical education. However, there are two challenges when applying the approach to assess LLM therapists at scale. Ethically, asking humans to frequently mimic clients and exposing them to potentially harmful LLM outputs can be risky and unsafe. Technically, it can be difficult to consistently compare the performances of different LLM therapists interacting with the same client. To this end, we adopt LLMs to simulate clients and propose ClientCAST, a client-centered approach to assessing LLM therapists by client simulation. Specifically, the simulated client is utilized to interact with LLM therapists and complete questionnaires related to the interaction. Based on the questionnaire results, we assess LLM therapists from three client-centered aspects: session outcome, therapeutic alliance, and self-reported feelings. We conduct experiments to examine the reliability of ClientCAST and use it to evaluate LLMs therapists implemented by Claude-3, GPT-3.5, LLaMA3-70B, and Mixtral 8*7B. Codes are released at https://github.com/wangjs9/ClientCAST.
Forward citations
Cited by 6 Pith papers
-
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
Dialogue agents aligned via DPO on preference pairs mined from simulated conversations improve engagement scores against the same simulator, with smaller and partially inconsistent human evaluation evidence.
-
"Is This Really a Human Peer Supporter?": Misalignments Between Peer Supporters and Experts in LLM-Supported Interactions
Mixed-methods studies of an LLM-supported peer support system uncover systematic misalignments where mental health experts flag critical safety and fidelity issues in peer responses that the supporters themselves do n...
-
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
The DynToM benchmark shows ten LLMs average 33.0% accuracy versus 77.7% for humans, and models lose the most accuracy on questions about mental-state changes across scenarios.
-
StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation
A multi-LLM agent workflow with questionnaire-to-story grounding and dynamic MI code control generates steerable motivational-interviewing dialogues and improves measured strategy adherence across six models.
-
"I Said Things I Needed to Hear Myself": Peer Support as an Emotional, Organisational, and Sociotechnical Practice in Singapore
Volunteer peer supporters in Singapore experience emotional labour, organisational gaps, and ambivalence toward AI, yielding design implications for human-centred support technologies.
-
Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning
An open mental-well-being chatbot demo with retrieval-augmented responses, synthetic dialogue generation, and agentic self-care planning reports a small human study favoring its retrieval-augmented version.
Discussion (0). Continue with ORCID to comment.