Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Empirical Modeling of Therapist-Client Dynamics in Psychotherapy Using LLM-Based Assessments

T0 review · 4 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Using LLM ratings of 1,610 therapy sessions, this paper claims that moment-to-moment therapist empathy and exploration directly increase client disclosure and shape emotional expression, while rapport works as a contextual moderator rather

desk verdict Real measurement contribution, but the moderation headline is not in the model table and the emotion PCA section contradicts itself; worth a serious referee after major revision. read the letter →

arxiv 2602.12450 v2 pith:COLM7XB6 submitted 2026-02-12 cs.CY

classification cs.CY
keywords psychotherapyprocessLLM-basedassessmentstructuralequationmodelingtherapistempathytherapeuticrapportself-disclosurenegativeemotionexpressionturn-levelanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that core psychotherapy processes can be measured automatically and modeled moment-to-moment from transcripts. It claims that therapist empathy and exploration in one turn directly increase client self-disclosure and shift emotional expression in the next turn, while rapport built in prior sessions does not directly amplify disclosure or emotions but changes how those behaviors matter. If right, this would make therapy process research scalable and give trainers, supervisors, and AI-based tools a concrete, turn-level signal to work from.

What carries the argument

The carrying mechanism is a two-step pipeline: (1) construct-specific LLM prompts that rate each therapist turn, client turn, and session segment on theory-grounded scales (EPITOME empathy, WAI bond items, intimacy-based disclosure, PANAS-like emotions), with human validation (mean Pearson r=.66); (2) a structural equation model with utterance-level variables (therapist behavior in prior turn → client outcome in current turn), session-level rapport from prior session, and controls for prior client state and session index. PCA collapses nine emotions into self-directed negative, outward-directed negative, and surprise factors.

What would settle it

Re-estimate the SEM using only emotion items with adequate LLM–human agreement (anger, contempt, disgust, enjoyment, surprise) or replace the LLM emotion scores with human-coded labels on a subset of sessions; if the empathy coefficient for self-directed negative emotion collapses toward zero, the asymmetry claim fails. Also test the interaction term rapport×empathy in the model; if it is not significant, the moderator claim is unsupported.

Watch

Extended reading notes

Core claim

Across 1,610 sessions and about 243,000 utterances, the authors built LLM-based scores for therapist empathy (emotional reaction, interpretation, exploration, reflection), rapport (observer-rated bond), and client self-disclosure and nine emotions, validated them against human raters, and fitted a structural equation model. They found empathy (β=0.12) and exploration (β=0.10) predict self-directed negative emotions, empathy weakly predicts outward-directed emotions (β=0.03), both predict disclosure (β=0.09), and prior-session rapport predicts slightly lower self-directed negative emotion (β=-0.02) rather than more disclosure. The paper interprets this as a dual-route model: empathy and explo

Load-bearing premise

The claim that empathy mainly increases self-directed rather than outward-directed negative emotion depends on LLM scores for fear, anxiety, sadness, and depression being meaningful measures, but their agreement with human raters is low enough that the 95% confidence interval includes zero.

Editorial extensions

If this is right

  • If correct, turn-level process modeling becomes feasible for thousands of sessions without human coding, enabling large-scale tests of therapy mechanisms.
  • Therapist training systems could give moment-by-moment feedback—'this empathic reflection is likely to increase disclosure'—grounded in the estimated paths.
  • The dual-route model predicts that empathy's effect on internal distress expression is larger than its effect on outward anger, a testable signature of supported emotional processing.
  • Designers of chatbot therapists could embed these contingencies to make simulated clients respond to trainee empathy and exploration in realistic ways.
  • Observer-rated rapport not directly increasing disclosure would push process research to distinguish client-perceived alliance from observer-rated bond.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The low LLM–human agreement for fear, anxiety, sadness, and depression (ICCs 0.45–0.53 with confidence intervals including zero) means the self-directed negative emotion factor—and therefore the empathy-asymmetry result—may rest on noisy measurements; a focused re-analysis with human-coded emotion labels would test this.
  • The paper claims rapport moderates associations, but the reported SEM shows only direct effects; an explicit empathy-by-rapport interaction term would be the natural falsifier.
  • Because rapport is observer-rated (WAI-O) rather than client-rated, the null direct effect may reflect who is doing the rating; client-rated alliance could behave differently.
  • The Alexander Street corpus is a single, professionally transcribed, Western collection; generalizing to other languages, modalities, or cultural display norms remains untested.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper develops an LLM-based measurement pipeline for psychotherapy process constructs—therapist empathy (Emotional Reaction, Interpretation, Exploration, Reflection), self-disclosure, nine emotions, and session-level rapport—and validates the scores against human annotations. It then applies the pipeline to roughly 243k utterances from 1,610 Alexander Street sessions and fits a multivariate SEM with utterance-level therapist behaviors and session-level rapport as predictors of next-turn client disclosure and emotional expression. The reported main effects suggest empathy and exploration predict disclosure and emotional expression, and prior-session rapport is negatively associated with self-directed negative emotion. The abstract and conclusion additionally claim that rapport moderates the effects of therapist empathy and exploration on client affect.

Significance. If the measurement-validation results hold, the framework is a useful contribution: the authors provide detailed prompt rubrics, a human-validation protocol with ICC reporting, and a large transcript corpus. This could support scalable psychotherapy process research. However, as reported, the SEM does not estimate the moderation interactions that the headline claim requires, and the low LLM–human reliability of the internalizing emotion indicators threatens the most distinctive emotion findings. The paper would need substantial reanalysis—not just copyediting—to support its central claims.

major comments (4)
  1. [§5.1, Table 7, §7, Abstract] The abstract and §7 claim that rapport is a contextual moderator that 'dampens empathy’s association with negative emotions and conditions the impact of exploration.' However, Table 7 contains only main effects (log session ID, Rapport, Exploration, Empathy); no Rapport×Empathy or Rapport×Exploration interaction is defined or estimated, and Figure 3 shows only direct arrows. §6.4 itself states that 'future work should test sequential mediation and moderation more directly,' conceding that moderation was not actually tested. This is not a wording issue: the distinguishing part of the headline claim is not identified by the reported analysis.
  2. [§4.4.2, Table 5, §5] The self-directed negative emotion factor is built from fear, anxiety, sadness, and depression plus reverse enjoyment. In Table 5, the LLM–human ICC 95% confidence intervals for these four internalizing emotions all include zero (Fear 0.454 [-0.13,0.71]; Anxiety 0.460 [-0.19,0.74]; Sadness 0.526 [-0.16,0.77]; Depression 0.500 [-0.06,0.72]). The key emotion findings—empathy’s larger association with self-directed than outward-directed negative emotion (β=0.12 vs 0.03) and rapport’s negative association (β=-0.02)—depend on this composite. No sensitivity analysis excluding or downweighting these low-ICC indicators is reported. The acknowledgment in §4.4.2 does not address the impact on the SEM estimates.
  3. [Abstract, §5.1, Table 7] The abstract states that the SEM estimates moment-to-moment relationships 'controlling for prior client state and context,' but the predictor set in Table 7 includes only log session ID, Rapport, Exploration, and Empathy. There is no lagged client outcome (e.g., previous-turn disclosure or previous-turn emotion) in the model. Without such a control, the reported associations may be confounded by client-level trends and autocorrelation, and the causal language in §7 ('exert immediate effects') is not justified.
  4. [§5, PCA] The manuscript gives two contradictory PCA descriptions for the emotion items. The first says 'Across nine client emotion indicators (N=122,939), PCA produced a three-component structure explaining 71.5%' with surprise as the third component. The following paragraph says 'PCA yielded a two-component solution that explained about 59% of the variance' and that surprise 'did not relate meaningfully to either component (uniqueness U2=.97).' Since Table 7 and Figure 3 treat surprise as a separate outcome, the three-component solution appears to be the one used, but the text must be reconciled. This inconsistency undermines the reproducibility of the factor construction.
minor comments (5)
  1. [Figure 3 caption] The caption says 'rapport was associated with reduced negative emotions and a slight decrease in disclosure.' Table 7 shows the disclosure coefficient as 0.00 (p=.71), which is not a decrease. Please correct the caption.
  2. [§4.4.1, Table 5] Describing an ICC of 0.45 as 'fair' is misleading when the 95% confidence interval includes zero. Report the CI explicitly in the text and avoid implying acceptable reliability.
  3. [Appendix A.2] Typo in the user prompt: 'herapist’s Response' should be 'therapist’s Response.'
  4. [General] The paper does not include a data or code availability statement. Given the detailed prompts and the reproducibility value of the pipeline, such a statement would strengthen the manuscript.
  5. [Abstract] The abstract reports a mean Pearson r=.66, but the validation section uses mixed metrics (e.g., F1 for Reflection, ICC and r for others). Please clarify how this average is computed and over which constructs.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation; a minor self-citation and an overclaimed moderation statement do not make the SEM results circular.

full rationale

The paper's empirical chain is not circular. The LLM measures are validated against human annotations on a stratified sample of the Alexander Street corpus (Table 5), independently of the SEM estimates; the SEM paths in Table 7 are then estimated from those validated scores. The PCA factor definitions in Section 5 are data-driven on the analysis corpus, but they are not fitted to the SEM outcomes and do not coincide with the reported regression coefficients, so this is standard two-step estimation, not an input recycled as an output. The only self-citation of note is Table 4's rapport calibration on 7 Cups WAI-O-S data [92], which shares a co-author (R. Kraut); however, Table 5 reports an independent LLM-human ICC = 0.806 for rapport on the Alexander Street sample, so the calibration citation is not load-bearing. Separately, the Conclusion's claim that rapport 'moderates' empathy/exploration effects (Section 7) goes beyond the fitted model: Table 7 contains only main effects, and Section 6.4 concedes that 'future work should test sequential mediation and moderation more directly.' This is an inference/overclaim issue, not a circularity, so it does not raise the circularity score. The internalizing-emotion ICCs (Table 5) are a measurement-validity concern but likewise not a circular input-output identity.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The central empirical claims rest on (i) LLM scores being valid for internalizing emotions despite ICC CIs crossing zero, (ii) factor structures extracted from the same corpus used to test hypotheses, (iii) observer-rated rapport standing in for client experience, and (iv) a temporal ordering that supports causal interpretation without lagged outcomes.

free parameters (3)
  • PCA component solution for client emotions = 3 components: self-directed negative (sadness .45, anxiety .43, fear .40, depression .42, enjoyment -.33), outward-direc
    The outcome variables in the SEM are not fixed a priori; they are factor scores from PCA run on the same 122,939 emotion observations used to test the hypotheses. The contradictory two- vs three-component descriptions in Section 5 make the choice ambiguous.
  • PCA component solution for therapist empathy = 2 components: general empathy (interpretation .62, reflection .57, reaction .49), exploration (.94)
    Defines the empathy and exploration predictors from the same corpus used in SEM, rather than from a pre-registered or external factor structure.
  • LLM prompt rubrics and context-window choices = 0–2/0–1/1–5/1–7 scales; 2 prior utterances for empathy/disclosure; 5 prior utterances for emotion; 1 segment for rapport
    The measurement layer is tuned by iterative prompt refinement against human annotations; the chosen rubric scales and context windows are researcher decisions that shape every downstream coefficient.
assumptions (6)
  • domain assumption LLM-generated emotion scores for fear, anxiety, sadness, and depression are reliable enough to serve as indicators of a self-directed negative emotion factor.
    Entered into PCA/SEM despite ICC confidence intervals that include zero (Table 5: Fear ICC 0.454 [-0.13,0.71], Anxiety 0.460 [-0.19,0.74], Depression 0.500 [-0.06,0.72], Sadness 0.526 [-0.16,0.77]).
  • domain assumption The Alexander Street corpus is an accurate, anonymized, representative sample of one-on-one psychotherapy.
    Section 3 states transcripts are professionally transcribed and anonymized, but the corpus is a convenience/licensed sample (37 clients) with primarily Western contexts, as the limitations section acknowledges.
  • ad hoc to paper The PCA labels 'self-directed negative emotions' and 'outward-directed negative emotions' correspond to clinically meaningful dimensions.
    The labels are interpretive names attached to components extracted from the same data; no external validity evidence is provided for these specific factors.
  • domain assumption Treating Unique Session ID as a nested effect sufficiently accounts for the non-independence of utterances within clients and sessions.
    Section 5 says lavaan treats Unique Session ID as a nested effect, but there are only 37 clients and no therapist identifiers, so client-level and therapist-level clustering may be misspecified.
  • domain assumption Observer-rated WAI-O rapport captures the client's experience of the therapeutic bond.
    The paper itself notes in Section 6.1 that observer, therapist, and client ratings are not interchangeable, yet SEM uses observer-rated rapport to test rapport hypotheses.
  • domain assumption The temporal ordering of therapist turn -> client turn, combined with the included covariates, supports the causal language ('directly shaped').
    No lagged client state is included in Table 7, and no therapist/client random effects, so omitted-variable confounding remains.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Empirical Modeling of Therapist-Client Dynamics in Psychotherapy Using LLM-Based Assessments." pith.science (2026). https://pith.science/paper/COLM7XB6

@misc{pith2026260212450,
  author       = {Pith},
  title        = {Pith review of: Empirical Modeling of Therapist-Client Dynamics in Psychotherapy Using LLM-Based Assessments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/COLM7XB6}},
  note         = {Machine review of arXiv:2602.12450}
}
read the original abstract

Psychotherapy is a primary treatment for many mental health conditions, yet the interplay among therapist behaviors, client responses, and the therapeutic relationship is difficult to study at scale, as process research has relied on labor-intensive human coding. We develop and validate a computational framework for modeling therapist-client interaction, using large language models (LLMs) to measure therapist behaviors (empathy, exploration), relational quality (rapport), and client outcomes (self-disclosure, self-directed and outward-directed negative emotion). After validating model-generated scores against human annotations (ICC = 0.45-0.81; rapport 0.81, self-disclosure 0.78), we apply these measures to roughly 2,000 hours of transcripts from the Alexander Street corpus and use Structural Equation Modeling to estimate moment-to-moment relationships among therapist behaviors, rapport, and subsequent client responses, controlling for prior client state and context. Therapist empathy and exploration directly predict increased client disclosure and shifts in emotional expression; empathy is more strongly associated with self-directed than outward-directed negative emotion, suggesting greater acknowledgment of internal distress, while exploration increases disclosure and emotional elaboration. Rapport does not directly amplify disclosure or emotional intensity but instead moderates the associations between therapist behaviors and client affect, potentially contributing to reductions in internal distress. These results show that LLM-based measurement combined with structural modeling can capture core therapeutic processes at scale, with empathy and exploration acting directly and rapport as a contextual moderator, providing a foundation for precision modeling of psychotherapy and for scalable therapist training and AI-supported clinical education.

Figures

Figures reproduced from arXiv: 2602.12450 by the authors.

Figure 1
Figure 1. Hypothesized Relationships Hypothesized links among therapist behaviors, the therapist–client relationship, and [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The automatic assessment process starts with designing LLM prompts to generate quantitative psychological assess [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. The results show that rapport was associated with reduced negative emotions and a slight decrease in disclosure. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Principal Component Analysis (PCA) of eight [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 5
Figure 5. Figure 5: Correlation matrix of emotions E Rapport Trend [PITH_FULL_IMAGE:figures/full_fig_p019_5.png]
Figure 6
Figure 6. Figure 6: Rapport trend across first 30 sessions. The trend shows rapport doesn’t change much after initial sessions. [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Empirically Grounded Adaptive Virtual Patient for Psychotherapy Training: Disclosure That Responds to Therapist Micro-Skills

    cs.CY 2026-06 unverdicted novelty 5.0 of 10

    An adaptive virtual patient uses a structural equation model fitted to nearly 2000 hours of real transcripts to dynamically update disclosure in response to therapist empathy and exploration, showing adaptation in a s...

Reference graph

Works this paper leans on

134 extracted references · 28 canonical work pages · cited by 1 Pith paper

  1. [1]

    Alaa Ali Abd-Alrazaq, Asma Rababeh, Mohannad Alajlani, Bridgette M Bewick, and Mowafa Househ. 2020. Effectiveness and safety of using chatbots to improve mental health: systematic review and meta-analysis.Journal of medical Internet research22, 7 (2020), e16021

  2. [2]

    Alexander Street. 2025. Counseling and Psychotherapy Transcripts, Client Narratives, and Reference Works. Subscription Database. Accessed on 2025- 04-28. URL typically provided via institutional access. General product infor- mation: https://alexanderstreet.com/products/counseling-and-psychotherapy- transcripts-client-narratives-and-reference-works

  3. [3]

    Tim Althoff, Kevin Clark, and Jure Leskovec. 2016. Large-scale analysis of counseling conversations: An application of natural language processing to mental health.Transactions of the Association for Computational Linguistics4 (2016), 463–476

  4. [4]

    Mario Alvarez-Jimenez, Simon Rice, Simon D’Alfonso, Steven Leicester, Sarah Bendall, Ingrid Pryor, Penni Russon, Carla McEnery, Olga Santesteban-Echarri, Gustavo Da Costa, et al. 2020. A novel multimodal digital service (moderated online social therapy+) for help-seeking young people experiencing mental ill-health: pilot evaluation within a national youth...

  5. [5]

    Hill, and Dennis M

    Morgan Anvari, Clara E. Hill, and Dennis M. Kivlighan. 2020. Therapist skills associated with client emotional expression in psychodynamic psychotherapy. Chen et al. Psychotherapy Research30, 7 (2020), 900–911. doi:10.1080/10503307.2019.1680901

  6. [6]

    Anvari, Vardaan Dua, Jose Lima-Rosas, Clara E

    Morgan S. Anvari, Vardaan Dua, Jose Lima-Rosas, Clara E. Hill, and Dennis M. Kivlighan. 2022. Facilitating exploration in psychodynamic psychotherapy: Therapist skills and client attachment style.Journal of Counseling Psychology 69, 3 (2022), 348–360. doi:10.1037/cou0000582

  7. [7]

    JinYeong Bak, Chin-Yew Lin, and Alice Oh. 2014. Self-disclosure topic model for classifying and analyzing Twitter conversations. InProceedings of the 2014 Con- ference on Empirical Methods in Natural Language Processing (EMNLP). Associa- tion for Computational Linguistics, Doha, Qatar, 1986–1996. doi:10.3115/v1/D14- 1213

  8. [8]

    Sairam Balani and Munmun De Choudhury. 2015. Detecting and characterizing mental health related self-disclosure in social media. InProceedings of the 33rd annual ACM conference extended abstracts on human factors in computing systems. 1373–1378

Show all 134 references
  1. [9]

    Luke Balcombe and Diego De Leo. 2022. Human-Computer Interaction in Digital Mental Health.Informatics9, 1 (2022), 14. doi:10.3390/informatics9010014

  2. [10]

    Scott A Baldwin, Bruce E Wampold, and Zac E Imel. 2007. Untangling the alliance-outcome correlation: exploring the relative importance of therapist and patient variability in the alliance.Journal of consulting and clinical psychology 75, 6 (2007), 842

  3. [11]

    Lambert, and David Saxon

    Michael Barkham, Wolfgang Lutz, Michael J. Lambert, and David Saxon. 2017. Therapist effects, effective therapists, and the law of variability. InHow and Why Are Some Therapists Better than Others? Understanding Therapist Effects, Louis G. Castonguay and Clara E. Hill (Eds.). ...

  4. [12]

    Chiara Berardi, Marta Antonini, Zoe Jordan, et al. 2024. Barriers and Facilitators to the Implementation of Digital Technologies in Mental Health Systems: A Qualitative Systematic Review to Inform a Policy Framework.BMC Health Services Research24, 1 (2024), 243. doi:10.1186/s1...

  5. [13]

    Eleanor R Burgess, Sean A Munson, David C Mohr, and Madhu C Reddy

  6. [14]

    Arnow, Robert Kraut, and Diyi Yang

    Alicja Chaszczewicz, Raj Sanjay Shah, Ryan Louie, Bruce A. Arnow, Robert Kraut, and Diyi Yang. 2024. Multi-Level Feedback Generation with Large Language Models for Empowering Novice Peer Counselors. arXiv:2403.15482 [cs.CL] https://arxiv.org/abs/2403.15482

  7. [15]

    Xiang Cheng, Chengyan Pan, Minjun Zhao, Deyang Li, Fangchao Liu, Xinyu Zhang, Xiao Zhang, and Yong Liu. 2025. Revisiting Chain-of-Thought Prompt- ing: Zero-shot Can Be Stronger than Few-shot. arXiv:2506.14641 [cs.CL] https://arxiv.org/abs/2506.14641

  8. [16]

    Prerna Chikersal, Danielle Belgrave, Gavin Doherty, Angel Enrique, Jorge E Palacios, Derek Richards, and Anja Thieme. 2020. Understanding client support strategies to improve clinical outcomes in an online mental health intervention. InProceedings of the 2020 CHI conference on...

  9. [17]

    Cicchetti

    Domenic V. Cicchetti. 1994. Guidelines, criteria, and rules of thumb for evaluat- ing normed and standardized assessment instruments in psychology.Psycho- logical Assessment6, 4 (1994), 284–290. doi:10.1037/1040-3590.6.4.284

  10. [18]

    Constantino, Amy E

    Michael J. Constantino, Amy E. Coyne, James F. Boswell, and Adam Visla. 2019. Promoting treatment credibility: Therapist contributions. InPsychotherapy Relationships That Work, Volume 1: Evidence-Based Therapist Contributions(3 ed.), John C. Norcross and Michael J. Lambert (Ed...

  11. [19]

    David A Cook, Joshua Overgaard, V Shane Pankratz, Guilherme Del Fiol, and Chris A Aakre. 2025. Virtual patients using large language models: Scalable, contextualized simulation of clinician-patient dialogue with feedback.Journal of Medical Internet Research27 (2025), e68486

  12. [20]

    David Coyle and Gavin Doherty. 2009. Clinical evaluations and collaborative design: developing new technologies for mental healthcare interventions. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 2051–2060

  13. [21]

    John R Crawford and Julie D Henry. 2004. The Positive and Negative Affect Schedule (PANAS): Construct validity, measurement properties and normative data in a large non-clinical sample.British journal of clinical psychology43, 3 (2004), 245–265

  14. [22]

    Katherine Gibbons

    Paul Crits-Christoph and M. Katherine Gibbons. 2021. The process–outcome research paradigm: Past accomplishments and present challenges. InBergin and Garfield’s Handbook of Psychotherapy and Behavior Change(50th anniversary edition ed.), Michael Barkham, Wolfgang Lutz, and Lou...

  15. [23]

    Simon D’Alfonso, Reeva Lederman, Sandra Bucci, and Katherine Berry. 2020. The Digital Therapeutic Alliance and Human-Computer Interaction.JMIR Mental Health7, 12 (2020), e21895. doi:10.2196/21895

  16. [24]

    Eaton, Norman Abeles, and Mary J

    Timothy T. Eaton, Norman Abeles, and Mary J. Gutfreund. 1988. Therapeutic alliance and outcome: Impact of treatment length and pretreatment symptoma- tology.Psychotherapy: Theory, Research, Practice, Training25, 4 (1988), 536–542. doi:10.1037/h0085379

  17. [25]

    Paul Ekman, Tim Dalgleish, and M Power. 1999. Basic emotions.San Francisco, USA(1999)

  18. [26]

    Bohart, Jeanne C

    Robert Elliott, Arthur C. Bohart, Jeanne C. Watson, and Leslie S. Greenberg

  19. [27]

    Rachel Elvins and Jonathan Green. 2008. The conceptualization and measure- ment of therapeutic alliance: An empirical review.Clinical psychology review 28, 7 (2008), 1167–1187

  20. [28]

    2012.Emotion regulation and culture: The effects of cultural models of self on Western and East Asian differences in suppression and reappraisal

    Joshua Stephen Eng. 2012.Emotion regulation and culture: The effects of cultural models of self on Western and East Asian differences in suppression and reappraisal. University of California, Berkeley

  21. [29]

    Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. 2023. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. arXiv:2309.16797 [cs.CL] https://arxiv.org/abs/2309. 16797

  22. [30]

    Del Re, Bruce E

    Christoph Flückiger, Anthony C. Del Re, Bruce E. Wampold, and Adam O. Horvath. 2018. The Alliance in Adult Psychotherapy: A Meta-Analytic Synthesis. Psychotherapy55, 4 (2018), 316–340. doi:10.1037/pst0000172

  23. [31]

    Ellie Fossey, Carol Harvey, Mohammadreza R Mokhtari, and Graham N Mead- ows. 2012. Self-rated assessment of needs for mental health care: a qualitative analysis.Community mental health journal48, 4 (2012), 407–419

  24. [32]

    Vasudha Gidugu, E Sally Rogers, Steven Harrington, Mihoko Maru, Gene John- son, Julie Cohee, and Jennifer Hinkel. 2015. Individual Peer Support: A Qualita- tive Study of Mechanisms of Its Effectiveness. 445–452 pages

  25. [33]

    Leslie Greenberg. 2008. Emotion and Cognition in Psychotherapy: The Trans- forming Power of Affect.Canadian Psychology/Psychologie Canadienne49, 1 (2008), 49–59. doi:10.1037/0708-5591.49.1.49

  26. [34]

    Greenberg

    Leslie S. Greenberg. 2007. A guide to conducting a task analysis of psychother- apeutic change.Psychotherapy Research17, 1 (2007), 15–30. doi:10.1080/ 10503300600720390

  27. [35]

    Smith, and Sachin Kumar

    Carolina Hatanpää, Noah A. Smith, and Sachin Kumar. 2025. On Distributional Robustness of In-Context Learning for Text Classification. InSecond Workshop on Test-Time Adaptation: Putting Updates to the Test! at ICML 2025. https: //openreview.net/forum?id=fJgOCX5rDp

  28. [36]

    K. J. Heaton, C. E. Hill, D. E. Krause, and R. J. Clift. 1995. Comparing molar and molecular methods of judging therapist techniques.Psychotherapy Research5, 2 (1995), 141–153. doi:10.1080/10503309512331331266

  29. [37]

    2021.Essentials of Conversation Analysis

    Alexa Hepburn and Jonathan Potter. 2021.Essentials of Conversation Analysis. American Psychological Association, Washington, DC

  30. [39]

    Hill, Janet E

    Clara E. Hill, Janet E. Helms, Veronica Tichenor, Scott B. Spiegel, Kevin E. O’Grady, and Earl S. Perry. 1988. Effects of therapist response modes in brief psychotherapy.Journal of Counseling Psychology35, 3 (1988), 222–233. doi:10. 1037/0022-0167.35.3.222

  31. [40]

    Hill and Ian S

    Clara E. Hill and Ian S. Kellems. 2002. Development and use of the Helping Skills Measure to assess client perceptions of the effects of training and of helping skills in sessions.Journal of Counseling Psychology49, 2 (2002), 264–272. doi:10.1037/0022-0167.49.2.264

  32. [41]

    Hill and Sarah Knox

    Clara E. Hill and Sarah Knox. 2002. Self-disclosure. InPsychotherapy re- lationships that work: Therapist contributions and responsiveness to patients, John C. Norcross (Ed.). Oxford University Press, 255–265. doi:10.1093/med: psych/9780195142187.003.0011

  33. [42]

    C. E. Hill and S. Knox. 2021.Essentials of Consensual Qualitative Research. American Psychological Association, Washington, DC

  34. [43]

    Hill, John C

    Clara E. Hill, John C. Norcross, and the Steering Committee. 2023. Introduction to Psychotherapy Skills and Methods That Work. InPsychotherapy Skills and Methods That Work, Clara E. Hill and John C. Norcross (Eds.). Oxford University Press, Chapter 1. doi:10.1093/oso/978019761...

  35. [44]

    Adam O Horvath and Leslie S Greenberg. 1989. Development and validation of the Working Alliance Inventory.Journal of counseling psychology36, 2 (1989), 223

  36. [45]

    Horvath and B

    Adam O. Horvath and B. Dianne Symonds. 1991. Relation between Working Alliance and Outcome in Psychotherapy: A Meta-Analysis.Journal of Counseling Psychology38, 2 (1991), 139–149. doi:10.1037/0022-0167.38.2.139

  37. [46]

    Annika Howells, Itai Ivtzan, and Francisco Jose Eiroa-Orosa. 2016. Putting the ‘app’in happiness: a randomised controlled trial of a smartphone-based mindfulness intervention to enhance wellbeing.Journal of happiness studies17, 1 (2016), 163–185

  38. [47]

    Hoyt, Janet N

    William T. Hoyt, Janet N. Melby, and Sharlene A. Wolchik. 1994. Structural equation modeling of client change in therapy.Journal of Consulting and Clinical Psychology62, 2 (1994), 297–305. doi:10.1037/0022-006X.62.2.297

  39. [48]

    Shang-Ling Hsu, Raj Sanjay Shah, Prathik Senthil, Zahra Ashktorab, Casey Dugan, Werner Geyer, and Diyi Yang. 2025. Helping the Helper: Supporting Peer Counselors via AI-Empowered Practice and Feedback.Proceedings of the ACM on Human-Computer Interaction9, CSCW2, Article CSCW09...

  40. [49]

    Daniel E Jimenez, Mijung Park, Daniel Rosen, Jin hui Joo, David Martinez Garza, Elliott R Weinstein, Kyaien Conner, Caroline Silva, and Olivia Okereke. 2022. Centering culture in mental health: differences in diagnosis, treatment, and access to care among older people of color...

  41. [50]

    Jolliffe

    Ian T. Jolliffe. 2002.Principal Component Analysis(2nd ed.). Springer

  42. [51]

    Kahn and Ann M

    Jeffrey H. Kahn and Ann M. Garrison. 2009. Emotional self-disclosure and emo- tional avoidance: Relations with symptoms of depression and anxiety.Journal of Counseling Psychology56, 4 (2009), 573–584. doi:10.1037/a0016574

  43. [52]

    Parsons, Jonathan Gratch, Anton Leuski, and Albert A

    Patrick Kenny, Thomas D. Parsons, Jonathan Gratch, Anton Leuski, and Albert A. Rizzo. 2007. Virtual Patients for Clinical Therapist Skills Training. InIntelligent Virtual Agents: 7th International Conference, IV A 2007, Proceedings (Lecture Notes in Computer Science, Vol. 4722...

  44. [53]

    Patrick G Kenny, Thomas D Parsons, AI Geer, ID ORCiD, and Thomas D Parsons

  45. [54]

    Heejung S Kim, David K Sherman, and Shelley E Taylor. 2008. Culture and social support.American psychologist63, 6 (2008), 518

  46. [55]

    M. Kim, J. Hong, and M. Ban. 2021. Mediating effects of emotional self-disclosure on the relationship between depression and quality of life for women undergo- ing in-vitro fertilization.International Journal of Environmental Research and Public Health18, 12 (2021), 6247. doi:...

  47. [56]

    1988.Rethinking Psychiatry: From Cultural Category to Per- sonal Experience

    Arthur Kleinman. 1988.Rethinking Psychiatry: From Cultural Category to Per- sonal Experience. Free Press

  48. [57]

    Rex B. Kline. 2015.Principles and Practice of Structural Equation Modeling(4th ed.). Guilford Press

  49. [58]

    S. Knox, S. Hess, D. Petersen, and C. E. Hill. 1997. A qualitative analysis of client perceptions of the effects of helpful therapist self-disclosure in long-term therapy.Journal of Counseling Psychology44, 3 (1997), 274–283. doi:10.1037/ 0022-0167.44.3.274

  50. [59]

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023. Large Language Models are Zero-Shot Reasoners. arXiv:2205.11916 [cs.CL] https://arxiv.org/abs/2205.11916

  51. [60]

    Koo and Mae Y

    Terry K. Koo and Mae Y. Li. 2016. A Guideline of Selecting and Reporting Intra- class Correlation Coefficients for Reliability Research.Journal of chiropractic medicine15, 2 (2016), 155–163

  52. [61]

    Spitzer, and Janet B

    Kurt Kroenke, Robert L. Spitzer, and Janet B. W. Williams. 2001. The PHQ-9: Validity of a brief depression severity measure.Journal of General Internal Medicine16, 9 (2001), 606–613. doi:10.1046/j.1525-1497.2001.016009606.x

  53. [62]

    Claire Lane and Stephen Rollnick. 2007. The use of simulated patients and role-play in communication skills training: a review of the literature to August 2005.Patient education and counseling67, 1-2 (2007), 13–20

  54. [63]

    Laska, Alan S

    Kevin M. Laska, Alan S. Gurman, and Bruce E. Wampold. 2014. Expanding the Lens of Evidence-Based Practice in Psychotherapy: A Common Factors Perspective.Psychotherapy51, 4 (2014), 467–481. doi:10.1037/a0034332

  55. [64]

    Matthew J Leach. 2005. Rapport: A key to treatment success.Complementary therapies in clinical practice11, 4 (2005), 262–265

  56. [65]

    Matthew J. Leach. 2005. Rapport: A Key to Treatment Success.Complementary Therapies in Clinical Practice11, 4 (2005), 262–265. doi:10.1016/j.ctcp.2005.05.005

  57. [66]

    Han Li, Renwen Zhang, Yi-Chieh Lee, Robert E Kraut, and David C Mohr. 2023. Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being.NPJ Digital Medicine6, 1 (2023), 236

  58. [67]

    Nangyeon Lim. 2016. Cultural differences in emotion: differences in emotional arousal level between the East and the West.Integrative medicine research5, 2 (2016), 105–109

  59. [68]

    Fan Liu, Wenshuo Chao, Naiqiang Tan, and Hao Liu. 2025. Bag of Tricks for Inference-time Computation of LLM Reasoning. arXiv:2502.07191 [cs.AI] https://arxiv.org/abs/2502.07191

  60. [69]

    Wolfgang Lutz et al . 2021. Measuring, predicting, and tracking change in psychotherapy. InBergin and Garfield’s Handbook of Psychotherapy and Behavior Change(50th anniversary edition ed.), Michael Barkham, Wolfgang Lutz, and Louis G. Castonguay (Eds.). Wiley, Hoboken, NJ, 89–134

  61. [70]

    Hazel Rose Markus and Shinobu Kitayama. 1991. Culture and the self: Implica- tions for cognition, emotion, and motivation.Psychological Review98, 2 (1991), 224–253. doi:10.1037/0033-295X.98.2.224

  62. [71]

    Martin, John P

    Daniel J. Martin, John P. Garske, and M. Katherine Davis. 2000. Relation of the therapeutic alliance with outcome and other variables: A meta-analytic review.Journal of Consulting and Clinical Psychology68, 3 (2000), 438–450. doi:10.1037/0022-006X.68.3.438

  63. [72]

    McCarthy and Jacques P

    Kevin S. McCarthy and Jacques P. Barber. 2009. The Multitheoretical List of Therapeutic Interventions (MULTI): Initial report.Psychotherapy Research19, 1 (2009), 96–113. doi:10.1080/10503300802524343

  64. [73]

    McGraw and S

    Kenneth O. McGraw and S. P. Wong. 1996. Forming inferences about some intraclass correlation coefficients.Psychological Methods1, 1 (1996), 30–46. doi:10.1037/1082-989X.1.1.30

  65. [74]

    Scott D Miller, BL Duncan, Jeb Brown, JA Sparks, and DA Claud. 2003. The out- come rating scale: A preliminary study of the reliability, validity, and feasibility of a brief visual analog measure.Journal of brief Therapy2, 2 (2003), 91–100

  66. [75]

    Miller and Stephen Rollnick

    William R. Miller and Stephen Rollnick. 2023.Motivational interviewing: helping people change and grow(fourth edition ed.). The Guilford Press, New York London

  67. [76]

    Do June Min, Verónica Pérez-Rosas, Kenneth Resnicow, and Rada Mihalcea. 2022. PAIR: Prompt-Aware margIn Ranking for Counselor Reflection Scoring in Moti- vational Interviewing. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Associatio...

  68. [77]

    Mohammad

    Saif M. Mohammad. 2016. Sentiment analysis: Detecting valence, emotions, and other affectual states from text. InEmotion Measurement. Elsevier, 201–237. doi:10.48550/arXiv.2005.11882

  69. [78]

    Bengt Muthén and Linda K. Muthén. 2002. Beyond SEM: General latent variable modeling.Behaviormetrika29, 1 (2002), 81–117. doi:10.2333/bhmk.29.81

  70. [79]

    Norcross and Michael J

    John C. Norcross and Michael J. Lambert (Eds.). 2019.Psychotherapy Rela- tionships That Work, Volume 1: Evidence-Based Therapist Contributions(3 ed.). Oxford University Press, New York

  71. [80]

    Norcross and Michael J

    John C. Norcross and Michael J. Lambert. 2019.Psychotherapy Relationships That Work: Volume 1: Evidence-Based Therapist Contributions(3rd ed.). Oxford University Press, New York, NY. https://global.oup.com/academic/product/ psychotherapy-relationships-that-work-9780190843953

  72. [81]

    Shubham Parashar, Blake Olson, Sambhav Khurana, Eric Li, Hongyi Ling, James Caverlee, and Shuiwang Ji. 2025. Inference-Time Computations for LLM Rea- soning and Planning: A Benchmark and Insights. arXiv:2502.12521 [cs.AI] https://arxiv.org/abs/2502.12521

  73. [82]

    The Only Way Out Is Through

    Antonio Pascual-Leone and Leslie S. Greenberg. 2007. Emotional Processing in Experiential Therapy: Why “The Only Way Out Is Through”.Journal of Consulting and Clinical Psychology75, 6 (2007), 875–887. doi:10.1037/0022- 006X.75.6.875

  74. [83]

    Greenberg, and Juan Pascual-Leone

    Antonio Pascual-Leone, Leslie S. Greenberg, and Juan Pascual-Leone. 2009. Developments in task analysis: New methods to study change.Psychotherapy Research19, 4-5 (2009), 527–542. doi:10.1080/10503300902897797

  75. [84]

    Aderonke Bamgbose Pederson. 2023. Management of depression in black people: effects of cultural issues.Psychiatric annals53, 3 (2023), 122–125

  76. [85]

    Paul R Peluso and Robert R Freund. 2018. Therapist and client emotional expression and psychotherapy outcomes: A meta-analysis.Psychotherapy55, 4 (2018), 461

  77. [86]

    Peluso and Robert R

    Paul R. Peluso and Robert R. Freund. 2018. Therapist and Client Emotional Expression and Psychotherapy Outcomes: A Meta-Analysis.Psychotherapy55, 4 (2018), 461–472. doi:10.1037/pst0000165

  78. [87]

    Hill, and Dennis M

    Megan Prass, Arcadia Ewell, Clara E. Hill, and Dennis M. Kivlighan. 2021. Solicited and unsolicited therapist advice in psychodynamic psychotherapy: Is it advised?Counselling Psychology Quarterly34, 2 (2021), 253–274. doi:10.1080/ 09515070.2020.1723492

  79. [88]

    StatPearls Publishing. 2023. Psychotherapy and Therapeutic Relationship. In StatPearls. StatPearls Publishing, Treasure Island (FL). https://www.ncbi.nlm. nih.gov/books/NBK608012/

  80. [89]

    Philip Resnik, Aaron Garron, and Rebecca Resnik. 2015. Using topic modeling to improve prediction of neuroticism and depression in social media. InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics,...

  81. [90]

    Fujiko Robledo Yamamoto, Amy Voida, and Stephen Voida. 2021. From Therapy to Teletherapy: Relocating Mental Health Services Online.Proceedings of the ACM on Human-Computer Interaction5, CSCW2, Article 364 (Oct. 2021), 364:1– 364:30 pages. doi:10.1145/3479508

  82. [91]

    Yves Rosseel. 2012. lavaan: An R Package for Structural Equation Modeling. Journal of Statistical Software48, 2 (2012), 1–36. doi:10.18637/jss.v048.i02

  83. [92]

    Kraut, and Diyi Yang

    Raj Sanjay Shah, Faye Holt, Shirley Anugrah Hayati, Aastha Agarwal, Yi- Chia Wang, Robert E. Kraut, and Diyi Yang. 2022. Modeling Motivational Interviewing Strategies on an Online Peer-to-Peer Counseling Platform.Pro- ceedings of the ACM on Human-Computer Interaction6, CSCW2 (...

  84. [93]

    Lin, Adam S

    Ashish Sharma, Inna W. Lin, Adam S. Miner, David C. Atkins, and Tim Althoff

  85. [94]

    Miner, David C

    Ashish Sharma, Adam S. Miner, David C. Atkins, and Tim Althoff. 2020. A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support. (2020). arXiv:2009.08441 https://doi.org/10.48550/ ARXIV.2009.08441

  86. [95]

    Petr Slovak and Sean A Munson. 2024. Hci contributions in mental health: A modular framework to guide psychosocial intervention design. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–21

  87. [96]

    Ewan Soubutts, Pranita Shrestha, Brittany I Davidson, Chengcheng Qu, Char- lotte Mindel, Aaron Sefi, Paul Marshall, and Roisin McNaney. 2024. Challenges and opportunities for the design of inclusive digital mental health tools: under- standing culturally diverse young people’s...

  88. [97]

    Spitzer, Kurt Kroenke, Janet B

    Robert L. Spitzer, Kurt Kroenke, Janet B. W. Williams, and Bernd Löwe. 2006. A brief measure for assessing generalized anxiety disorder: the GAD-7.Archives of Internal Medicine166, 10 (2006), 1092–1097. doi:10.1001/archinte.166.10.1092

  89. [98]

    William B. Stiles. 1979. Verbal response modes and psychotherapeutic technique. Psychiatry42 (1979), 49–62

  90. [99]

    William B. Stiles. 1988. Psychotherapy process–outcome correlations may be misleading.Psychotherapy25, 1 (1988), 27–35. doi:10.1037/h0085320

  91. [100]

    2016.Counseling the Culturally Diverse: Theory and Practice

    Derald Wing Sue and David Sue. 2016.Counseling the Culturally Diverse: Theory and Practice. John Wiley & Sons

  92. [101]

    Tausczik and James W

    Yla R. Tausczik and James W. Pennebaker. 2010. The psychological meaning of words: LIWC and computerized text analysis methods.Journal of Language and Social Psychology29, 1 (2010), 24–54. doi:10.1177/0261927X09351676

  93. [102]

    Anja Thieme, Danielle Belgrave, and Gavin Doherty. 2020. Machine learning in mental health: A systematic review of the HCI literature to support the development of effective and implementable ML systems.ACM Transactions on Computer-Human Interaction (TOCHI)27, 5 (2020), 1–53

  94. [103]

    Anja Thieme, Maryann Hanratty, Maria Lyons, Jorge Palacios, Rita Faia Mar- ques, Cecily Morrison, and Gavin Doherty. 2023. Designing human-centered AI for mental health: Developing clinically relevant applications for online CBT treatment.ACM Transactions on Computer-Human Int...

  95. [104]

    Victoria Tichenor and Clara E Hill. 1989. A comparison of six measures of working alliance.Psychotherapy: Theory, Research, Practice, Training26, 2 (1989), 195

  96. [105]

    Victoria Tichenor and Clara E. Hill. 1989. A Comparison of Six Measures of Working Alliance.Psychotherapy: Theory, Research, Practice, Training26, 2 (1989), 195–199. doi:10.1037/h0085419

  97. [106]

    Town, Gillian E

    Joel M. Town, Gillian E. Hardy, Leigh McCullough, and Chris Stride. 2012. Patient affect experiencing following therapist interventions in short-term dynamic psychotherapy.Psychotherapy Research22, 2 (2012), 208–219. doi:10. 1080/10503307.2011.637243

  98. [107]

    Vaidyam, Haley Wisniewski, John D

    Aditya N. Vaidyam, Haley Wisniewski, John D. Halamka, Matcheri S. Kashavan, and John B. Torous. 2019. Chatbots and Conversational Agents in Mental Health: A Review of the Psychiatric Landscape.Canadian Journal of Psychiatry. Revue Canadienne de Psychiatrie64, 7 (2019), 456–464...

  99. [108]

    Girard, Lauren M

    Alexandria Vail, Jeffrey M. Girard, Lauren M. Bylsma, Jeffrey F. Cohn, Jay C. Fournier, Holly A. Swartz, and Louis-Philippe Morency. 2022. Toward Causal Understanding of Therapist-Client Relationships: A Study of Language Modality and Social Entrainment. InProceedings of the 2...

  100. [109]

    Vondracek and Fred W

    Sarah I. Vondracek and Fred W. Vondracek. 1971. THE MANIPULATION AND MEASUREMENT OF SELF-DISCLOSURE IN PREADOLESCENTS.Merrill- Palmer Quarterly of Behavior and Development17, 1 (1971), 51–58. http://www. jstor.org/stable/23083609

  101. [110]

    Bruce E. Wampold. 2015. How Important Are the Common Factors in Psy- chotherapy? An Update.World Psychiatry14, 3 (2015), 270–277. doi:10.1002/ wps.20238

  102. [111]

    Zixiu Wu, Simone Balloccu, Vivek Kumar, Rim Helaoui, Ehud Reiter, Diego Re- forgiato Recupero, and Daniele Riboni. 2022. Anno-MI: A Dataset of Expert- Annotated Counselling Dialogues. InICASSP 2022 - IEEE International Confer- ence on Acoustics, Speech and Signal Processing. I...

  103. [112]

    Wenjie Yang, Anna Fang, Raj Sanjay Shah, Yash Mathur, Diyi Yang, Haiyi Zhu, and Robert E Kraut. 2024. What makes digital support effective? how therapeutic skills affect clinical well-being.Proceedings of the ACM on Human-Computer Interaction8, CSCW1 (2024), 1–29

  104. [113]

    Jianwen Zeng, Wenhao Qi, Shiying Shen, Xin Liu, Sixie Li, Bing Wang, Chaoqun Dong, Xiaohong Zhu, Yankai Shi, Xiajing Lou, et al. 2025. Embracing the Future of Medical Education With Large Language Model–Based Virtual Patients: Scoping Review.Journal of Medical Internet Researc...

  105. [114]

    You are a professional evaluator specializing in assessing therapist empathy. Your evaluations are based on clinical best practices, focusing on accuracy, nuance, and consistency

    Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2023. Large Language Models Are Human-Level Prompt Engineers. arXiv:2211.01910 [cs.LG] https://arxiv.org/abs/2211.01910 A Constructs Prompts A.1 Empathy - EPITOME System Rol...

  106. [119]

    The weather has been nice lately

    G (General, No Disclosure) - The client talks about external topics without sharing personal experiences, thoughts, or emotions. - Example: "The weather has been nice lately."

  107. [120]

    I started a new job last week

    M (Medium Disclosure) - The client shares non-sensitive personal experiences, plans, or general life updates. This includes age, occupation and hobbies, personal events, past history, future plans, or details about family members. - Example: "I started a new job last week."

  108. [121]

    Session Background: {summary} Client:

    H (High Disclosure) - The client reveals deeply personal, sensitive, or emotional experiences, including thoughts about mental health, relationships, or private struggles. This includes personal characteristics, problematic behaviors, physical appearance concerns, or wishful t...

  109. [122]

    Instructions:

    Depression Use the following 5-point Likert scale to rate the intensity of each emotion: 1 - Very slightly or not at all, 2 - A little, 3 - Moderately, 4 - Quite a bit, 5 - Extremely. Instructions:

  110. [123]

    Carefully analyze central utterance within the context of the preceding conversation

  111. [124]

    For each of the nine emotions, identify specific conversational cues or statements that support your rating

  112. [125]

    If no clear emotion is present, assign a rating of 1 to all emotions

    Provide a rating (1-5) for each emotion. If no clear emotion is present, assign a rating of 1 to all emotions

  113. [126]

    "" A.5 Rapport System Role Message Chen et al

    Provide a evidence-based rationale for your rating, referencing contextual cues from the previous utterances. Avoid making assumptions not supported by the dialogue. Contextual Analysis: Use the utterances that precede the central utterance as the context (could be empty): {co...

  114. [127]

    There is a mutual liking between the client and therapist

  115. [128]

    The client feels confident in the therapist's ability to help the client

  116. [129]

    The client feels that the therapist appreciates him/her as a person

  117. [130]

    Very strong evidence against

    There is mutual trust between the client and therapist. Use the following 7-point Likert scale for each aspect: 1: "Very strong evidence against", 2: "Considerable evidence against", 3: "Some evidence against", 4: "No evidence", 5: "Some evidence", 6: "Considerable evidence", ...

  118. [131]

    Carefully read through the conversation log provided below

  119. [132]

    For each of the four bond aspects, identify specific conversational cues or statements that support your rating

  120. [133]

    Provide a rating (1-7) for each aspect and include a brief explanation, citing specific parts of the conversation that influenced your decision

  121. [134]

    After rating each aspect, give an **overall bond rating** (1-7) for the entire conversation based on a holistic view of the interaction

  122. [135]

    Avoid making assumptions not supported by the dialogue

    Your analysis should be objective, relying only on observable cues from the text. Avoid making assumptions not supported by the dialogue. Now, based on four key aspects of bond, rate the provided conversation log on a 7-point Likert scale (1-7),considering mutual liking, confi...

  123. [2018]

    Psychotherapy55, 4 (2018), 399–410

    Therapist Empathy and Client Outcome: An Updated Meta-Analysis. Psychotherapy55, 4 (2018), 399–410. doi:10.1037/pst0000175

  124. [2023]

    doi:10.1038/s42256-022-00593-2

    Human–AI collaboration enables more empathic conversations in text- based peer-to-peer mental health support.Nature Machine Intelligence5, 1 (2023), 46–57. doi:10.1038/s42256-022-00593-2

  125. [2024]

    Virtual standardized LLM-AI patients for clinical practice.Annual Review of Cybertherapy And Telemedicine 2024(2024), 177

  126. [2025]

    InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems

    What’s In Your Kit? Mental Health Technology Kits for Depression Self-Management. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–19

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.