Pith. sign in

REVIEW 5 major objections 4 minor 1 cited by

AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing

T0 review · 5 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A SHAP-derived chain-of-thought prompt raises GPT-4's predicted user-perceived quality in motivational interviewing from 38.45% to 46.84%, closing most of the gap to human therapists at 52.69%.

desk verdict A useful prompt-engineering pipeline that is undermined by an overclaimed title and a self-referential evaluation loop. read the letter →

arxiv 2505.17380 v1 pith:CPK5CJCY submitted 2025-05-23 cs.CL

classification cs.CL
keywords motivationalinterviewinguser-perceivedqualityGPT-4chain-of-thoughtpromptingSHAPexplainableAILIWCmentalhealthLLMevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a general-purpose LLM can be guided toward therapist-level performance in motivational interviewing (MI) by building an explainable, user-centered quality metric and converting its findings into a prompt. The authors train a model to predict whether a response would be perceived as high quality, identify 17 MI-consistent and MI-inconsistent verbal behaviors that drive that prediction, and embed the resulting rules in a chain-of-thought prompt for GPT-4. With that prompt, GPT-4's share of high-quality responses rises from 38.45% to 46.84%, while human therapists remain at 52.69%. The paper argues this is evidence that LLMs can approach, though not yet match, human therapists in core MI skills, and that the framework provides a transparent path for improving mental-health chatbots.

What carries the argument

The load-bearing machinery is a UPQ-centered computational evaluation framework. UPQ (user-perceived quality) is the extrinsic binary outcome ('high' or 'low'); the intrinsic inputs are 17 SHAP-selected behavioral metrics spanning MI-consistent behaviors (simple and complex reflection), MI-inconsistent behaviors (unsolicited advice, warnings, directive 'need' language, raising concerns without permission), and other linguistic cues (empathy perception words, apostrophe use, analytic thinking markers, language style matching). An SVM classifier with LIWC features predicts UPQ, SHAP values rank and direct each feature, and the framework then converts the ranked direction of influence into a customized zero-shot chain-of-thought prompt that tells GPT-4 which behaviors to increase and which to decrease.

What would settle it

Blind-rate the actual human, GPT-4, and GPT-4-Prompted transcripts with MITI-trained coders or real clients and check whether their judgments reproduce the paper's ordering (human 52.69% above prompted 46.84% above baseline 38.45%); if human raters see no improvement from prompting, or place prompted GPT-4 at or above the human level, the central claim is settled by the data.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that user-perceived quality (UPQ) in MI dialogues can be predicted from linguistic cues and MI strategy frequencies, and that SHAP-based explanations of that predictor reveal actionable behavioral targets. Applying those targets as a customized zero-shot chain-of-thought prompt reduces GPT-4's unsolicited advice (Cohen's d = -0.258 relative to baseline) and raises simple reflections, empathy-related perception words, and work-related discussion. In the headline comparison, prompted GPT-4 reaches 46.84% high-UPQ responses versus 38.45% for unprompted GPT-4 ($\chi^2 = 0.292$, $p = 0.002$, OR = 1.21) and remains below human therapists at 52.69% ($\chi^2 = 4.966$, $p = 0.026$, OR = 0.81). The paper reads this as GPT-4 achieving performance comparable to human therapists in core MI skills while remaining marginally inferior overall.

Load-bearing premise

The load-bearing premise is that the SVM's predicted 'user-perceived quality' equals how real clients would experience the responses, even though part of the training data was generated and labelled by GPT-4; if the proxy is not a faithful stand-in for user perception, the comparison to human therapists collapses.

Editorial extensions

If this is right

  • If the framework's UPQ proxy is trusted, the same measure-explain-prompt loop can be applied to other LLMs in MI without retraining the model, giving an inexpensive route to safer mental-health chatbots.
  • The 17-metric list gives developers and regulators a concrete behavioral checklist: increase reflections, empathy perception words, and permission-seeking; reduce unsolicited advice, warnings, directive 'need' language, and informal contractions.
  • Because prompted GPT-4 still trails human therapists on directive commands, warnings, unsolicited concerns, and session structure, further gains would require fine-tuning or multi-stage alignment rather than prompting alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit is to evaluate the same prompt with actual client ratings or MITI-trained expert judgments; if those human measures reproduce the SVM's ordering, the framework's usefulness is independently confirmed.
  • Because the UPQ proxy was trained partly on GPT-4-augmented labels, an inference beyond the reported results is that the true gap between prompted GPT-4 and human therapists could differ under human-only labels, and this should be checked before deployment.
  • The mechanism suggests a general recipe for other clinical communication skills, such as cognitive behavioral therapy fidelity: use explainable models to find behavior-level targets, then encode them as prompts or fine-tuning objectives, though the paper only demonstrates this for MI.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The manuscript proposes a computational framework for evaluating GPT-4's motivational interviewing (MI) performance. The authors use 133 human-annotated MI dialogues, augment them with GPT-4-generated dialogues to 598, train an SVM (with LIWC features) to predict a binary user-perceived quality (UPQ) label, apply SHAP to identify 17 intrinsic metrics, design a customized chain-of-thought prompt, and compare baseline GPT-4, prompted GPT-4, and human therapists on UPQ and linguistic metrics. The paper's central claim is that the prompted model achieves 'therapist-level' MI responses, while its own primary comparison shows a statistically significant shortfall.

Significance. If the UPQ proxy were externally validated, the SHAP-to-prompt pipeline would be a useful contribution to automated MI fidelity assessment and transparent prompt optimization. The paper also deserves credit for reporting its limitations and for providing a detailed comparison of intrinsic linguistic metrics. However, the main contribution as stated cannot be accepted: the primary endpoint contradicts the therapist-level claim, and the evaluation pipeline has a self-referential component that makes the reported improvements non-independent. The framework's significance is therefore contingent on new human validation, not on the present results.

major comments (5)
  1. [§IV.C.2, Table VIII] The central claim in the title, the abstract, and Contribution (1) is contradicted by the paper's primary endpoint. GPT-4-Prompted achieves high UPQ in 46.84% of responses versus 52.69% for human therapists, with McNemar χ²=4.966, p=0.026, OR=0.81, indicating a statistically significant deficit rather than equivalence or non-inferiority. No equivalence margin or non-inferiority test is reported, so 'comparable' and 'therapist-level' are not supported. Section V.C.1's statement that 'LLMs perform comparably to human therapists in MI' is also inconsistent with Section V.A.1's statement that GPT-4 'perform[s] significantly below human therapists.'
  2. [§III.A, §IV.A.2, §V.B] The evaluation is self-referential. Section III.A states that GPT-4 was used to expand the 133 human-annotated dialogues to 598; Section IV.A.2 then uses an SVM trained on this augmented set to assign UPQ labels to GPT-4 responses. If the augmented samples carry GPT-4's own linguistic regularities, the classifier can be expected to favor GPT-4 in the human-versus-GPT comparisons. Moreover, the SHAP features extracted from this same SVM are used to construct the prompt whose improvement is measured with the same SVM, so the improvement is not independent evidence of framework effectiveness. The authors' own limitation statement in Section V.B acknowledges that the automated evaluation 'fell short of comprehensively capturing interviewers' subjective experiences' and 'often failed to accurately reflect clients' authentic feedback or the MI quality.' An analysis restricted to the original 133 human-annotated dialogues, plus external human ratings of LLM outputs, would be needed to break this circularity.
  3. [§IV.A.2, Table III] The UPQ values attributed to GPT-4 are model predictions, not direct user ratings. The paper reports classification accuracy of 0.9703 on the augmented dataset, but no calibration or validation of the predicted UPQ against actual user-perceived quality for LLM-generated responses is provided. All between-group comparisons, including the headline 38.45% and 46.84% figures, therefore inherit an unquantified proxy bias, and the high accuracy on a dataset that was itself augmented by GPT-4 does not establish measurement validity.
  4. [§III.D, Table VII] The customized chain-of-thought prompt is not included in the manuscript; Table VII gives a prompt framework, but Section IV.C.1 refers to 'comprehensive details' in supplementary materials, which are not part of this submission. Because the entire validation phase rests on the effect of this prompt, the intervention cannot be reproduced or independently assessed without the full prompt text and the exact GPT-4 inference settings (temperature, decoding, number of runs).
  5. [Table VIII] The row for 'GPT-4-Prompted vs GPT-4' reports χ²=0.292, p=0.002, OR=1.21. A McNemar chi-square of 0.292 with one degree of freedom corresponds to p≈0.59, so the reported p-value appears to be inconsistent with the stated test statistic. Since this row is the statistical basis for the claim that prompting improves UPQ, the authors should correct the value or explain the discrepancy.
minor comments (4)
  1. [Throughout] There are numerous typographical errors and formatting issues, e.g., 'GPT-4-PRMOPTED' in the Table VIII caption, 'Motivational Interview in by intrinsic metrics' in the Section IV.B.2 heading, and inconsistent use of 'GPT4' versus 'GPT-4'.
  2. [Table I and §III.A] The dataset unit is unclear: Table I reports 'Utterance Rounds' with a maximum of 598, while the text describes 598 video dialogues; these should be reconciled because the statistical unit determines the validity of the paired tests.
  3. [Abstract and Table VIII] The abstract's '(OR=1.21)' for the human-versus-prompted comparison is inconsistent with Table VIII's OR=0.81 for the same comparison; the direction and coding of the odds ratio should be stated explicitly.
  4. [Table II, Table IV, §IV.B.1] The reported ROC AUC values are inconsistent: Table II reports 0.9923 for the LIWC SVM, Table IV reports 0.9898 for the 'Original SVM Model,' and Section IV.B.1 says the pruned model's AUC rose from 0.9898 to 0.9918; these numbers should be reconciled.

Circularity Check

2 steps flagged · score 6.0 of 10

The reported prompt-driven improvement in UPQ is read back from the same SVM whose SHAP features were used to write the prompt, and that SVM was trained partly on GPT-4-augmented dialogues; the predicted gain is therefore partially forced by construction.

  1. fitted input called prediction [Section IV.C, Tables VII-VIII]
    "As demonstrated in Table VII, this study aimed to enhance the response quality of GPT-4 in the MI task by designing a customized prompt based on the computational evaluation framework. ... a McNemar test was conducted ... revealed a statistically significant improvement in UPQ for GPT-4-Prompted responses compared to standard GPT-4 responses ... (46.84% vs. 38.45%, χ² = 0.292, p = 0.002)."

    The computational evaluation framework is the SHAP interpretation of the pruned SVM, and the customized prompt instructs GPT-4 to increase or decrease exactly the features the SVM weights (e.g., less 'advice', more 'Complex Reflection'). The same SVM then assigns the UPQ labels whose rates are compared in Table VIII. The reported gain from 38.45% to 46.84% is therefore a closed-loop scoring result: the input was perturbed in the direction of the fitted model's own decision function, and the outcome was read from that same fitted model. No independent human rating or held-out instrument measures the prompted responses. The paper itself concedes that the automated evaluation 'often failed to accurately reflect clients' authentic feedback or the MI quality.'

  2. other [Section III.A and Section IV.A.2]
    "The dataset originally contained only 133 video dialogues and exhibited significant class imbalance ... this study employed GPT-4 to expand the dataset, resulting in an augmented dataset with 598 video dialogues. ... The predictive model, an SVM classifier based on LIWC features, was utilized to assess the UPQ of responses generated by GPT-4."

    The measurement instrument for GPT-4's UPQ is an SVM trained on a corpus in which GPT-4 itself generated the majority of dialogues (598 augmented vs. 133 human-annotated originals). Scoring GPT-4's responses with that SVM means the 'external' comparison between GPT-4 and human therapists in Table III is partly an evaluation of GPT-4 by a model that learned from GPT-4's own text. The target system thus contributes to constructing the yardstick that labels the target system, making the comparison self-referential rather than an independent benchmark.

full rationale

The paper's central empirical claim that customized prompts improve GPT-4's MI performance is not established by an independent measurement. The improvement is read from the same pruned SVM that produced the SHAP feature importances used to design the prompt, so the 8.39-point increase in predicted high-UPQ is partially forced by construction: the prompt changes the classifier's inputs in the directions its own weights favor, and the classifier's outputs are then presented as evidence. This is compounded by the fact that the classifier was trained on a dataset augmented with GPT-4-generated dialogues, so the evaluation yardstick itself is not independent of the system being assessed. There is, however, some independent grounding: 133 human-annotated dialogues anchor the original UPQ labels, and the direction of the prompt (less unsolicited advice, more reflection) is consistent with MI theory and external literature. No load-bearing self-citation chain or imported uniqueness theorem appears in the paper. The title's 'therapist-level' wording is also not supported by the paper's own Table VIII (46.84% vs. 52.69%, p = 0.026), but that is an overstatement rather than a circularity. Overall, the prompt-improvement result is partially circular, yielding a score of 6.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central comparison depends on the UPQ classifier being a valid proxy for user perception and on GPT-4-augmented training data being representative. The paper reports no human validation of predicted UPQ for GPT-4 transcripts, no code, and no prompt text. The SHAP-selected features are used both to explain and to modify behavior, which creates a self-referential loop.

free parameters (3)
  • SVM hyperparameters = not reported
    Selected by GridSearchCV on the augmented training set; the classifier's outputs are used as UPQ predictions for GPT-4.
  • 17-feature subset = feature list in Table V
    The RFECV and SHAP forward-selection procedure chooses features by model performance on the same dataset, and the same rankings are used to design the customized prompt.
  • Customized zero-shot CoT prompt wording = not released
    The prompt is hand-constructed from SHAP findings; because the wording is absent, the reported improvements cannot be separated from prompt design choices.
assumptions (4)
  • standard math SHAP values for a trained SVM are valid feature attributions.
    Used in Section III.C for global and local interpretation; SHAP is approximate for kernel-based models but is treated as exact in the paper.
  • domain assumption LIWC and LSM features capture MI-relevant therapist behavior.
    Section III.B treats LIWC output and LSM as primary predictors of UPQ; no validation is given that these lexical proxies align with established MI fidelity coding such as MITI or MISC.
  • ad hoc to paper GPT-4-generated data augmentation preserves the original UPQ distribution and feature-label relationships.
    Section III.A expands 133 human dialogues to 598 by GPT-4 generation, but the paper does not report human verification of the augmented labels.
  • ad hoc to paper The SVM-predicted UPQ for GPT-4 is a valid measure of user-perceived quality.
    Sections IV.A.2 and IV.C.2 compare GPT-4 to humans on classifier predictions, with no direct human ratings of GPT-4 transcripts.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing." pith.science (2026). https://pith.science/paper/CPK5CJCY

@misc{pith2026250517380,
  author       = {Pith},
  title        = {Pith review of: AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CPK5CJCY}},
  note         = {Machine review of arXiv:2505.17380}
}
read the original abstract

Large language models (LLMs) like GPT-4 show potential for scaling motivational interviewing (MI) in addiction care, but require systematic evaluation of therapeutic capabilities. We present a computational framework assessing user-perceived quality (UPQ) through expected and unexpected MI behaviors. Analyzing human therapist and GPT-4 MI sessions via human-AI collaboration, we developed predictive models integrating deep learning and explainable AI to identify 17 MI-consistent (MICO) and MI-inconsistent (MIIN) behavioral metrics. A customized chain-of-thought prompt improved GPT-4's MI performance, reducing inappropriate advice while enhancing reflections and empathy. Although GPT-4 remained marginally inferior to therapists overall, it demonstrated superior advice management capabilities. The model achieved measurable quality improvements through prompt engineering, yet showed limitations in addressing complex emotional nuances. This framework establishes a pathway for optimizing LLM-based therapeutic tools through targeted behavioral metric analysis and human-AI co-evaluation. Findings highlight both the scalability potential and current constraints of LLMs in clinical communication applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Effect of Static vs. Conversational AI-Generated Messages on Colorectal Cancer Screening Intent: a Randomized Controlled Trial

    cs.CY 2025-07 conditional novelty 6.0 of 10

    In 915 unscreened adults, a single tailored GPT-4.1 message matched an MI chatbot on colorectal screening intent and beat expert materials for stool-test intent but not colonoscopy intent.

Reference graph

Works this paper leans on

152 extracted references · 61 canonical work pages · cited by 1 Pith paper

  1. [1]

    advice” feature in the “Advice (without permission)

    Extraction of intrinsic metrics: Identifying psychological linguistic cues that influence UPQ through explainable ML The initial feature set in this study consisted of 127 variabl es, which was subsequently reduced to 21 through feature selection with RFECV. In this section, the global and local SHAP methods were utilized for sensitivity analysis and feat...

  2. [2]

    Comparison Analysis of UPQ Between GPT -4 and Human Therapists with Clients The predictive model, an SVM classifier based on LIWC features, was utilized to assess the UPQ of responses generated by GPT -4. Given that the UPQ of human therapists was represented as a binary variable, the McNemar test was employed for statistical analysis, alongside visualiza...

  3. [3]

    Advice (without permission)

    Evaluating GPT-4’s performance Motivational Interview in by intrinsic metrics To examine the impact of intrinsic metrics on UPQ, paired - sample t-tests were employed to compare the deployment of MI strategies and linguistic cues between human therapists and GPT-4, thereby elucidating differences in their performance across these intrinsic metrics. The re...

  4. [4]

    - advice Avoid advice without permission to boost autonomy, reduce resistance, and strengthen the alliance

    Developing zero-shot prompts: Utilizing insights from UPQ- centered explanatory modeling to enhance GPT-4’s responses Table VII PROMPT FRAMEWORK FOR ENHANCING GPT-4’S UPQ IN MOTIVATIONAL INTERVIEW Dimens ions of MI Categories of Intrinsic Metric Strategies promoting UPQ Strategies prompting Intrinsic Metrics Customized prompts Increase Usage Decrease Usag...

  5. [5]

    The results, presented in Fig

    Comparative analysis of the UPQ between GPT-4-Prompted and Human Therapists in MI To validate the effectiveness of the computational evaluation framework in enhancing GPT-4’s response performance on the extrinsic metric, a McNemar t est was conducted to examine whether significant differences in UPQ existed across three groups of respons es, with OR as th...

  6. [6]

    Advice (without permission)

    Comparison Analysis in the Intrinsic Metric between GPT - 4-Prompted and Human Therapists To assess whether customized prompts enhance GPT -4’s performance on intrinsic metrics, a paired -sample t -test was conducted to compare the use of intrinsic metrics between GPT- 4-Prompted and standard GPT-4 responses, as well as between GPT-4-Prompted responses an...

  7. [7]

    GPT-4's Capabilities in MIs Compared to Human Experts LLMs hold significant potential to enhance access to mental health support through scalable interventions capable of reaching extensive populations [34], [22] . To illustrate this potential, developers and end -users have shared anecdotal evidence on social media and other platforms, suggesting that LL...

  8. [8]

    double -edged sword

    The Role of the UPQ -Centered Evaluation Framework in Assessing GPT -4’s Performance in the Processes and Outcomes of Mis This study categorizes the intrinsic metrics into three primary types—MIIN behaviors, MICO behaviors, and other behaviors[105] —as introduced in the previous sections. MIIN behaviors refer to actions that conflict with the core princip...

Show all 152 references
  1. [9]

    The framework proposed in this study offers a promising preliminary exploration toward achieving this objective

    The Facilitating Role of the Integrative computational Evaluation Framework in Prompting LLM Performance in Mis As LLMs become increasingly integrated into both novel and existing mental health interventions —spanning commercial sectors[12], [14] and academic environments[13],...

  2. [10]

    Theoretical Significance From the theore tical significance, this study developed a computational evaluation framework to identify key differences in verbal behaviors between LLMs and human therapists during MI. This framework validated and extended Miller and Rollnick’s MI th...

  3. [11]

    Practical Significance From the practical significance, first, the study demonstrated that LLMs approached the MI performance of human therapists, highlighting their potential for mental health applications. As global demand for mental health support continues to exceed the av...

  4. [12]

    W. R. Miller and S. Rollnick, Motivational Interviewing: Helping People Change. Guilford Press, 2012

  5. [13]

    Language Models are Few -Shot Learners,

    T. B. Brown et al., “Language Models are Few -Shot Learners,” Jul. 22, 2020, arXiv: arXiv:2005.14165. doi: 10.48550/arXiv.2005.14165

  6. [14]

    GPT -4 Technical Report,

    OpenAI et al. , “GPT -4 Technical Report,” Mar. 04, 2024, arXiv: arXiv:2303.08774. doi: 10.48550/arXiv.2303.08774

  7. [15]

    Llama 2: Open Foundation and Fine -Tuned Chat Models,

    H. Touvron et al. , “Llama 2: Open Foundation and Fine -Tuned Chat Models,” Jul. 19, 2023, arXiv: arXiv:2307.09288. doi: 10.48550/arXiv.2307.09288

  8. [16]

    Bubeck et al

    S. Bubeck et al. , Sparks of A rtificial General Intelligence: Early experiments with GPT-4. 2023. doi: 10.48550/arXiv.2303.12712

  9. [17]

    Sun et al

    X. Sun et al. , “Eliciting Motivational Interviewing Skill Codes in Psychotherapy with LLMs: Joint 30th International Conference on Computational Linguis tics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024,” 2024 Joint International Co...

  10. [18]

    Artificial intelligence for global health,

    A. Hosny and H. J. W. L. Aerts, “Artificial intelligence for global health,” Science, vol. 366, no. 6468, pp. 955 –956, Nov. 2019, doi: 10.1126/science.aay5189

  11. [19]

    Mental Health and COVID -19: Early evidence of the pandemic’s impact: Scientific brief, 2 March 2022

    World Health Organization, “Mental Health and COVID -19: Early evidence of the pandemic’s impact: Scientific brief, 2 March 2022.” Accessed: Dec. 25, 2024. [Online]. Available: https://www.who.int/publications/i/item/WHO-2019-nCoV-Sci_Brief- Mental_health-2022.1

  12. [20]

    Building The Mental Health Workforce Capacity Needed To Treat Adults With Serious Mental Illnesses,

    M. Olfson, “Building The Mental Health Workforce Capacity Needed To Treat Adults With Serious Mental Illnesses,” Health Affairs, vol. 35, pp. 983–990, Jun. 2016, doi: 10.1377/hlthaff.2015.1619

  13. [21]

    Mental health stigma update: A review of consequences,

    A. Sickel, J. Seacat, and N. Nabors, “Mental health stigma update: A review of consequences,” Advances in Mental Health , vol. 12, pp. 202 – 215, Dec. 2014, doi: 10.1080/18374905.2014.11081898

  14. [22]

    Global Health Estimates 2021: Deaths by Cause, Age, Sex, by Country and by Region, 2000-2021

    G. World Health Organization, “Global Health Estimates 2021: Deaths by Cause, Age, Sex, by Country and by Region, 2000-2021.” Accessed: Dec. 06, 2024. [Online]. Available: https://www.who.int/data/gho/data/themes/mortality-and-global-health- estimates/ghe-leading-causes-of-death

  15. [23]

    Why generative AI (LLM) is ready for mental healthcare

    J. Hamilton, “Why generative AI (LLM) is ready for mental healthcare.” Accessed: Dec. 25, 2024. [Online]. Available: https://www.linkedin.com/pulse/why-generative-ai-chatgpt-ready- mental-healthcare-jose-hamilton-md

  16. [24]

    Large language models could change the future of behavioral healthcare: a proposal for responsible development and evaluation,

    E. Stade et al. , “Large language models could change the future of behavioral healthcare: a proposal for responsible development and evaluation,” npj Mental Health Research , vol. 3, Apr. 202 4, doi: 10.1038/s44184-024-00056-z

  17. [25]

    Youper: Artificial Intelligence For Mental Health Care

    Youper, “Youper: Artificial Intelligence For Mental Health Care.” Accessed: Dec. 25, 2024. [Online]. Available: https://www.youper.ai/

  18. [26]

    Sharma, A

    A. Sharma, A. Miner, D. Atkins, and T. Althoff, A Computational Approach to Understanding Empathy Expressed in Text -Based Mental Health Support. 2020. doi: 10.48550/arXiv.2009.08441

  19. [27]

    Sharma et al

    A. Sharma et al. , Cognitive Reframing of Negative Thoughts through Human-Language Model Interaction . 2023. doi: 10.48550/arXiv.2305.02466

  20. [28]

    IMBUE: Improving Interpersonal Effectiveness through Simulation and Just-in-time Feedback with Human-Language Model Interaction,

    I. W. Lin, A. Sharma, C. M. Rytting, A. S. Miner, J. Suh, and T. Althoff, “IMBUE: Improving Interpersonal Effectiveness through Simulation and Just-in-time Feedback with Human-Language Model Interaction,” Feb. 19, 2024, arXiv: arXiv:2402.12556. doi: 10.48550/arXiv.2402.12556

  21. [29]

    Sharma, K

    A. Sharma, K. Rushton, I. Lin, T. Nguyen, and T. Althoff, Facilitating Self-Guided Mental Health Interventions Through Human -Language Model Interaction: A Case Study of Cognitive Restructuring. 2024, p. 29. doi: 10.1145/3613904.3642761

  22. [30]

    OpenAI Employee Says She’s Never Tried Therapy But ChatGPT Is Pretty Much a Replacement For It

    N. Al -Sibai, “OpenAI Employee Says She’s Never Tried Therapy But ChatGPT Is Pretty Much a Replacement For It.” Accessed: Dec. 25, 2024. [Online]. Available: https://futurism.com/the -byte/openai-employee- chatgpt-therapy

  23. [31]

    Using ChatGPT as a therapist?,

    Cairo-TenThirteen, “Using ChatGPT as a therapist?,” r/ChatGPTPro. Accessed: Dec. 25, 2024. [Online]. Available: www.reddit.com/r/ChatGPTPro/comments/126rtvb/using_chatgpt_as_a_t herapist/

  24. [32]

    ChatGPT is better than my therapist, holy shit.,

    Mike2800, “ChatGPT is better than my therapist, holy shit.,” r/ChatGPT. Accessed: D ec. 25, 2024. [Online]. Available: www.reddit.com/r/ChatGPT/comments/zr5e17/chatgpt_is_better_than_m y_therapist_holy_shit/

  25. [33]

    Benefits and Harms of Large Language Models in Digital Mental Health,

    M. D. Choudhury, S. R. Pendse, and N. Kumar, “Benefits and Harms of Large Language Models in Digital Mental Health,” Nov. 07, 2023, arXiv: arXiv:2311.14693. doi: 10.48550/arXiv.2311.14693

  26. [34]

    The ChatGPT therapist will see you now: Navigating generative artificial intelligence’s potential in addiction medicine research and patient care,

    S. Tate, S. Fouladvand, J. H. Chen, and C. -Y. A. Chen, “The ChatGPT therapist will see you now: Navigating generative artificial intelligence’s potential in addiction medicine research and patient care,” Addiction, vol. 118, no. 12, pp. 2249–2251, Dec. 2023, doi: 10.1111/add.16341

  27. [35]

    Adapted large language models can outperform medical experts in clinical text summarization,

    D. Veen et al., “Adapted large language models can outperform medical experts in clinical text summarization,” Nature Medicine, vol. 30, pp. 1–9, Feb. 2024, doi: 10.1038/s41591-024-02855-5

  28. [36]

    A. T. Beck, Cognitive therapy and the emotional disorders . in Cognitive therapy and the emotional disorders. Oxford, England: International Universities Press, 1976, p. 356

  29. [37]

    Computer -Assisted Cog nitive-Behavior Therapy for Depression: A Systematic Review and Meta -Analysis,

    J. Wright et al. , “Computer -Assisted Cog nitive-Behavior Therapy for Depression: A Systematic Review and Meta -Analysis,” The Journal of Clinical Psychiatry, vol. 80, Mar. 2019, doi: 10.4088/JCP.18r12188

  30. [38]

    Comparing Automatic and Human Evaluation of NLG Systems,

    A. Belz and E. Reiter, “Comparing Automatic and Human Evaluation of NLG Systems,” in 11th Conference of the European Chapter of the Association for Computational Linguistics, D. McCarthy and S. Wintner, Eds., Trento, Italy: Association for Computational Linguistics, Apr. 2006,...

  31. [39]

    Why We Need New Evaluation Metrics for NLG,

    J. Novikova, O. Dušek, A. Cercas Curry, and V. Rieser, “Why We Need New Evaluation Metrics for NLG,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark: A ssociation for Computational Linguistics, 2017, pp. 2241 –

  32. [40]

    Evaluating a method of assessing competence in Motivational Interviewing: A study usin g simulated patients in the United Kingdom,

    G. A. Bennett, H. A. Roberts, T. E. Vaughan, J. A. Gibbins, and L. Rouse, “Evaluating a method of assessing competence in Motivational Interviewing: A study usin g simulated patients in the United Kingdom,” Addictive Behaviors , vol. 32, no. 1, pp. 69 –79, Jan. 2007, doi: 10.1...

  33. [41]

    J. L. Cochran and N. H. Cochran, The Heart of Counseling , 0 ed. Routledge, 2015. doi: 10.4324/9781315884066

  34. [42]

    Automated quality assessment of cognitive behavioral therapy sessions through highly contextualized language representations,

    N. Flemotomos, V. Martinez, Z. Chen, T. Creed, D. Atkins, and S. Narayanan, “Automated quality assessment of cognitive behavioral therapy sessions through highly contextualized language representations,” PLOS ONE , vol. 16, p. e0258639, Oct. 2021, doi: 10.1371/journal.pone.0258639

  35. [43]

    The Cognitive Therapy Scale: Psychometric properties.,

    T. M. Vallis, B. F. Shaw, and K. S. Dobson, “The Cognitive Therapy Scale: Psychometric properties.,” Journal of Consulting and Clinical Psychology, vol. 54, no. 3, pp. 381–385, Jun. 1986, doi: 10.1037/0022-006X.54.3.381

  36. [44]

    BLEU: a method for automatic evaluation of machine translation,

    K. Papineni, S. Roukos, T. Ward, and W. -J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting on Association for Computational Linguistics - ACL ’02, Philadelphia, Pennsylvania: Associatio n for Computational Lingu...

  37. [45]

    Re-evaluating the Role of Bleu in Machine Translation Research,

    C. Callison-Burch, M. Osborne, and P. Koehn, “Re-evaluating the Role of Bleu in Machine Translation Research,” in 11th Conference of the European Chapter of the Association f or Computational Linguistics , D. McCarthy and S. Wintner, Eds., Trento, Italy: Association for Comput...

  38. [46]

    Motivational Interviewing Skill Code (MISC) 2.1,

    J. Houck, “Motivational Interviewing Skill Code (MISC) 2.1,” vol. 2

  39. [47]

    Towards A Rigorous Science of Interpretable Machine Learning,

    F. Doshi -Velez and B. Kim , “Towards A Rigorous Science of Interpretable Machine Learning,” Mar. 02, 2017, arXiv: arXiv:1702.08608. doi: 10.48550/arXiv.1702.08608

  40. [48]

    Integrating explanation and prediction in computational social science,

    J. M. Hofman et al. , “Integrating explanation and prediction in computational social science,” Nature, vol. 595, no. 7866, pp. 181 –188, Jul. 2021, doi: 10.1038/s41586-021-03659-0

  41. [49]

    What is Motivational Interviewing?,

    S. Rollnick and W. R. Miller, “What is Motivational Interviewing?,” Behav. Cogn. Psychother. , vol. 23, no. 4, pp. 325 –334, Oct. 1995, doi: 10.1017/S135246580001643X

  42. [50]

    Motivational Interviewing,

    J. Hettema, J. Steele, and W. R. Miller, “Motivational Interviewing,” Annu. Rev. Clin. Psychol. , vol. 1, no. 1, pp. 91 –111, Apr. 2005, doi: 10.1146/annurev.clinpsy.1.102803.143833

  43. [51]

    Counsellors’ verbal behaviours and skills that elicit participants’ change or sustain talk in virtual motivational interviewing for physical activity among older adults,

    O. Akinrolie, S. Strachan, S. Webber, H. Chan, K. Messner, and R. Barclay, “Counsellors’ verbal behaviours and skills that elicit participants’ change or sustain talk in virtual motivational interviewing for physical activity among older adults,” Applied Psychology: Health and...

  44. [52]

    Counselor motivational interviewing skills and young adult change talk articulation during brief motivational interventions,

    J. Gaume, N. Bertholet, M. Faouzi, G. Gmel, and J. -B. Daeppen, “Counselor motivational interviewing skills and young adult change talk articulation during brief motivational interventions,” Journal of Substance Abuse Treatment , vol. 39, no. 3, pp. 272 –281, Oct. 2010, doi: 1...

  45. [53]

    Advancing Motivational Interviewing Training with Artificial Intelligence: ReadMI,

    P. J. Hershberger et al., “Advancing Motivational Interviewing Training with Artificial Intelligence: ReadMI,” AMEP, vol. Volume 12, pp. 613 – 618, Jun. 2021, doi: 10.2147/AMEP.S312373

  46. [54]

    The Motivational Interviewing Treatment Integrity Code (MITI 4): Rationale, Preliminary Reliability and Validity,

    T. B. Moyers, L. N. Rowell, J. K. Manuel, D. Ernst, and J. M. Houck, “The Motivational Interviewing Treatment Integrity Code (MITI 4): Rationale, Preliminary Reliability and Validity,” Journal of Substance Abuse Treatment, vol. 65, pp. 36–42, Jun. 2016, doi: 10.1016/j.jsat.2016.01.001

  47. [55]

    Lessons le arned from measuring fidelity with the Motivational Interviewing Treatment Integrity code (MITI 4),

    L. Kramer Schmidt, K. Andersen, A. S. Nielsen, and T. B. Moyers, “Lessons le arned from measuring fidelity with the Motivational Interviewing Treatment Integrity code (MITI 4),” Journal of Substance Abuse Treatment , vol. 97, pp. 59 –67, Feb. 2019, doi: 10.1016/j.jsat.2018.11.004

  48. [56]

    Adherence to the principl es of Motivational Interviewing, clients’ characteristics and behavior outcome in a smoking cessation and relapse prevention trial in women postpartum,

    J. R. Thyrian et al. , “Adherence to the principl es of Motivational Interviewing, clients’ characteristics and behavior outcome in a smoking cessation and relapse prevention trial in women postpartum,” Addictive Behaviors, vol. 32, no. 10, pp. 2297 –2303, Oct. 2007, doi: 10.1...

  49. [57]

    Can Motivational Interviewing in Emergency Care Reduce Alcohol Consumption in Young People? A Systematic Review and Meta -analysis,

    S. Kohler and A. Hofmann, “Can Motivational Interviewing in Emergency Care Reduce Alcohol Consumption in Young People? A Systematic Review and Meta -analysis,” Alcohol and Alcoholism, vol. 50, no. 2, pp. 107–117, Mar. 2015, doi: 10.1093/alcalc/agu098

  50. [58]

    Measuring client perceptions of motivational interviewing: factor analysis of the Client Evaluation of Motivational Interviewing scale,

    M. B. Madson et al. , “Measuring client perceptions of motivational interviewing: factor analysis of the Client Evaluation of Motivational Interviewing scale,” Journal of Substance Abuse Treatment, vol. 44, no. 3, pp. 330–335, Mar. 2013, doi: 10.1016/j.jsat.2012.08.015

  51. [59]

    Large Language Models in Mental Health Care: a Scoping Review,

    Y. Hua et al., “Large Language Models in Mental Health Care: a Scoping Review,” Aug. 21, 2024, arXiv: arXiv:2401.02984. doi: 10.48550/arXiv.2401.02984

  52. [60]

    The effectiveness and ineffectiveness of complex behavioral interventions: Impact of treatment fidelity,

    W. R. Miller and S. Rollnick, “The effectiveness and ineffectiveness of complex behavioral interventions: Impact of treatment fidelity,” Contemporary Clinical Trials, vol. 37, no. 2, pp. 234–241, Mar. 2014, doi: 10.1016/j.cct.2014.01.005

  53. [61]

    A Systematic Review of Psychometric Evaluation of Motivational Interviewing Integrity Measures,

    L. Wallace and F. Turner, “A Systematic Review of Psychometric Evaluation of Motivational Interviewing Integrity Measures,” Journal of Teaching in the Addictions , vol. 8, no. 1–2, pp. 84–123, Nov. 2009, doi: 10.1080/15332700903396655

  54. [62]

    Impact of Machine Learning in Natural Language Processing: A Review,

    T. P. Nagarhalli, V. Vaze, and N. K. Rana, “Impact of Machine Learning in Natural Language Processing: A Review,” in 2021 Third International Conference on Intelligent Communication Technolo gies and Virtual Mobile Networks (ICICV) , Feb. 2021, pp. 1529 –1534. doi: 10.1109/ICI...

  55. [63]

    Financial applications of machine learning: A literature review,

    N. Nazareth and Y. V. Ramana Reddy, “Financial applications of machine learning: A literature review,” Expert Systems with Applications, vol. 219, p. 119640, Jun. 2023, doi: 10.1016/j.eswa.2023.119640

  56. [64]

    Application progress of natural language processing technology in financial research,

    J. Xiao, J. Wang, W. Bao, T. Deng, and S. Bi, “Application progress of natural language processing technology in financial research,” Financial Engineering and Risk Management, vol. 7, no. 3, pp. 155–161, Jun. 2024, doi: 10.23977/ferm.2024.070320

  57. [65]

    Machine Learning in Education - a Survey of Current Research Trends,

    D. Kucak, V. Juricic, and G. Dambic, “Machine Learning in Education - a Survey of Current Research Trends,” in DAAAM Proceedings, 1st ed., vol. 1, B. Katalinic, Ed., DAAAM International Vienna, 2018, pp. 0406–0410. doi: 10.2507/29th.daaam.proceedings.059

  58. [66]

    Pathological Altruism - An Introduction,

    B. Oakley, A. Knafo-Noam, and M. Mcgrath, “Pathological Altruism - An Introduction,” Pathological Altruism , Dec. 2011, doi: 10.1093/acprof:oso/9780199738571.003.0014

  59. [67]

    Design feasibility of an automated, machine -learning based feedback system for motivational interviewing.,

    Z. E. Imel et al., “Design feasibility of an automated, machine -learning based feedback system for motivational interviewing.,” Psychotherapy, vol. 56, no. 2, pp. 318–328, Jun. 2019, doi: 10.1037/pst0000221

  60. [68]

    Attention is All you Need,

    A. Vaswani et al., “Attention is All you Need, ” in Advances in Neural Information Processing Systems, Curran Associates, Inc., 2017. Accessed: Dec. 10, 2024. [Online]. Available: https://papers.nips.cc/paper_files/paper/2017/hash/3f5ee243547dee91fbd 053c1c4a845aa-Abstract.html

  61. [69]

    Can Large Language Models Transform Computational Social Science?,

    C. Ziems, W. Held, O . Shaikh, J. Chen, Z. Zhang, and D. Yang, “Can Large Language Models Transform Computational Social Science?,” Computational Linguistics, vol. 50, no. 1, pp. 237 –291, Mar. 2024, doi: 10.1162/coli_a_00502

  62. [70]

    Language Models are Unsupervised Multitask Learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amode i, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” 2019. Accessed: Dec. 10, 2024. [Online]. Available: https://www.semanticscholar.org/paper/Language-Models-are- Unsupervised-Multitask-Learners-Radford- Wu...

  63. [71]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . 2018. doi: 10.48550/arXiv.1810.04805

  64. [72]

    Validity problems comparing values across cultures and possible solutions,

    K. Peng, R. E. Nisbett, and N. Y. C. Wong, “Validity problems comparing values across cultures and possible solutions,” Psychological Methods, vol. 2, no. 4, pp. 329–344, 1997, doi: 10.1037/1082-989X.2.4.329

  65. [73]

    Understanding the Benefits and Challenges of Using Large Language Model -based Conversational Agents for Mental Well-being Support,

    Z. Ma, Y. Mei, and Z. Su, “Understanding the Benefits and Challenges of Using Large Language Model -based Conversational Agents for Mental Well-being Support,” Ju l. 28, 2023, arXiv: arXiv:2307.15810. doi: 10.48550/arXiv.2307.15810

  66. [74]

    Assessing the Usability of a Chatbot for Mental Health Care,

    G. Cameron et al., “Assessing the Usability of a Chatbot for Mental Health Care,” 2019, pp. 121–132. doi: 10.1007/978-3-030-17705-8_11

  67. [75]

    Large Language Model for Mental Health: A Systematic Review,

    Z. Guo, A. Lai, J. H. Thygesen, J. Farrington, T. Keen, and K. Li, “Large Language Model for Mental Health: A Systematic Review,” Feb. 18, 2024. doi: 10.2196/preprints.57400

  68. [76]

    Farhat, ChatGPT as a Complementary Mental Health Resource: A Boon or a Bane

    F. Farhat, ChatGPT as a Complementary Mental Health Resource: A Boon or a Bane. 2023. doi: 10.20944/preprints202307.1479.v1

  69. [77]

    Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned,

    D. Ganguli et al. , “Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned,” Nov. 22, 2022, arXiv: arXiv:2209.07858. doi: 10.48550/arXiv.2209.07858

  70. [78]

    Factuality Challenges in the Era of Large Language Models,

    I. Augenstein et al., “Factuality Challenges in the Era of Large Language Models,” Oct. 10, 2023, arXiv: arXiv:2310.05189. doi: 10.48550/arXiv.2310.05189

  71. [79]

    Natural language processing of symptoms documented in free -text narratives of electronic health records: a systematic review,

    T. A. Koleck, C. Dreisbach, P. E. Bourne, and S. Bakken, “Natural language processing of symptoms documented in free -text narratives of electronic health records: a systematic review,” Journal of the American Medical Informatics Association, vol. 26, no. 4, pp. 364–379, Apr. ...

  72. [80]

    AI’s Role in Improving Social Connection and Oral Health for Older Adults: A Synergistic Approach,

    Q. X and W. B, “AI’s Role in Improving Social Connection and Oral Health for Older Adults: A Synergistic Approach,” JDR clinical and translational research , vol. 9, no. 3, Jul. 2024, doi: 10.1177/23800844231223097

  73. [81]

    Exploring the security and privacy risks of chatbots in messaging services: 22nd ACM Internet Measurement Conference, IMC 2022,

    J. Edu, C. Mulligan, F. Pierazzi, J. Polakis, G. Suarez-Tangil, and J. Such, “Exploring the security and privacy risks of chatbots in messaging services: 22nd ACM Internet Measurement Conference, IMC 2022,” IMC 2022 - Proceedings of the 2022 ACM Internet Measurement Conference...

  74. [82]

    Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks,

    M. A. Pimentel, C. Christophe, T. Raha, P. Munjal, P. K. Kanithi, and S. Khan, “Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks,” Jul. 29, 2024, arXiv: arXiv:2407.21072. doi: 10.48550/arXiv.2407.21072

  75. [83]

    A Survey on Evaluation of Large Language Models,

    Y. Chang et al., “A Survey on Evaluation of Large Language Models,” ACM Trans. Intell. Syst. Technol., vol. 15, no. 3, p. 39:1-39:45, 2024, doi: 10.1145/3641289

  76. [84]

    Instruction -Following Evaluation for Large Language Models,

    J. Zhou et al. , “Instruction -Following Evaluation for Large Language Models,” Nov. 14, 2023, arXiv: arXiv:2311.07911. doi: 10.48550/arXiv.2311.07911

  77. [85]

    Machine Learning, Natural Language Processing, and the Electronic Health Record: Innovations in Mental Health Services Research,

    J. B. Edgcomb and B. Zima, “Machine Learning, Natural Language Processing, and the Electronic Health Record: Innovations in Mental Health Services Research,” Psychiatric Services , vol. 70, p. appi.ps.2018004, Feb. 2019, doi: 10.1176/appi.ps.201800401

  78. [86]

    The Handbook of Computational Linguistics and Natural Language Processing,

    A. Clark and S. Lappin, “The Handbook of Computational Linguistics and Natural Language Processing,” 2010, pp. 197 –220. doi: 10.1002/9781444324044.ch8

  79. [87]

    Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge Generators,

    L. Chen et al., “Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge Generators,” Oct. 11, 2023, arXiv: arXiv:2310.07289. doi: 10.48550/arXiv.2310.07289

  80. [88]

    The emotional and mental health impact of the murder of George Floyd on the US population,

    J. C. Eichstaedt et al., “The emotional and mental health impact of the murder of George Floyd on the US population,” Proc. Natl. Acad. Sci. U.S.A., vol. 118, no. 39, p. e2109139118, Sep. 2021, doi: 10.1073/pnas.2109139118

  81. [89]

    Catatonia in autism and other neuro developmental disabilities: a state -of-the-art review,

    S. Moore, D. N. Amatya, M. M. Chu, and A. D. Besterman, “Catatonia in autism and other neuro developmental disabilities: a state -of-the-art review,” npj Mental Health Res , vol. 1, no. 1, p. 12, Sep. 2022, doi: 10.1038/s44184-022-00012-9

  82. [90]

    Facebook language predicts depression in medical records,

    J. C. Eichstaedt et al., “Facebook language predicts depression in medical records,” Proc. Natl. Acad. Sci. U.S.A., vol. 115, no. 44, pp. 11203–11208, Oct. 2018, doi: 10.1073/pnas.1802331115

  83. [91]

    The maintenance of behavioral change: The case for long - term follow -ups,

    R. M. Foxx, “The maintenance of behavioral change: The case for long - term follow -ups,” American Psychologist, vol. 68, no. 8, pp. 728 –736, 2013, doi: 10.1037/a0033713

  84. [92]

    REFRESH: Responsible and Efficient Feature Reselection guided by SHAP values,

    S. Sharma, S. Dutta, E. Albini, F. Lecue, D. Magazzeni, and M. Veloso, “REFRESH: Responsible and Efficient Feature Reselection guided by SHAP values,” in Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, in AIES ’23. New York, NY, USA : Association for Co...

  85. [93]

    Identification of suicidal behavior among psychiatrically hospitalized adolescents using natural language processing and machine learning of electronic health records,

    N. J. Carson et al. , “Identification of suicidal behavior among psychiatrically hospitalized adolescents using natural language processing and machine learning of electronic health records,” PLoS ONE, vol. 14, no. 2, p. e0211116, Feb. 2019, doi: 10.1371/journal.pone.0211116

  86. [94]

    Posttraumatic Stress Disorder Symptom Clusters and Perpetration of Intimate Partner Viole nce: Findings From a U.S. Nationally Representative Sample,

    K. Z. Smith, P. H. Smith, J. M. Violanti, P. T. Bartone, and G. G. Homish, “Posttraumatic Stress Disorder Symptom Clusters and Perpetration of Intimate Partner Viole nce: Findings From a U.S. Nationally Representative Sample,” Journal of Traumatic Stress , vol. 28, no. 5, pp. ...

  87. [95]

    Sentiment M easured in Hospital Discharge Notes Is Associated with Readmission and Mortality Risk: An Electronic Health Record Study,

    T. H. McCoy, V. M. Castro, A. Cagan, A. M. Roberson, I. S. Kohane, and R. H. Perlis, “Sentiment M easured in Hospital Discharge Notes Is Associated with Readmission and Mortality Risk: An Electronic Health Record Study,” PLoS ONE, vol. 10, no. 8, p. e0136341, Aug. 2015, doi: 1...

  88. [96]

    Towards Well-Being Measurement with Social Media Across Space, Time and Cultures: Three Generations of Progress

    O. Kjell Schwartz, “Towards Well-Being Measurement with Social Media Across Space, Time and Cultures: Three Generations of Progress.” Accessed: Dec. 29, 2024. [Online]. Available: https://worldhappiness.report/ed/2023/towards-well-being-measurement- with-social-media-across-sp...

  89. [97]

    Estimating geographic subjective well-being from Twitter: A comparison of dictionary and data-driven language methods,

    K. Jaidka, S. Giorgi, H. A. Schwartz, M. L. Kern, L. H. Ungar, and J. C. Eichstaedt, “Estimating geographic subjective well-being from Twitter: A comparison of dictionary and data-driven language methods,” Proc. Natl. Acad. Sci. U.S.A. , vol. 117, no. 19, pp. 10165 –10171, May...

  90. [98]

    An unexpected unity among methods for interpreting model predictions,

    S. Lundberg and S. -I. Lee, “An unexpected unity among methods for interpreting model predictions,” Nov. 2016, doi: 10.48550/arXiv.1611.07478

  91. [99]

    The past and current state of the Czech outpatient electronic prescription (eRecept),

    J. Bruthans, “The past and current state of the Czech outpatient electronic prescription (eRecept),” International Journal of Medical Informatics, vol. 123, pp. 49–53, Mar. 2019, doi: 10.1016/j.ijmedinf.2019.01.003

  92. [100]

    Grandparent(s) coresidence and physical activity/screen time among Latino children in the United States.,

    H. Xie, A. Ainsworth, and L. Caldwell, “Grandparent(s) coresidence and physical activity/screen time among Latino children in the United States.,” Families, Systems, & Health, vol. 39, no. 2, pp. 282–292, Jun. 2021, doi: 10.1037/fsh0000601

  93. [101]

    Significant and Distinctive n -Grams in Oncology Notes: A Text-Mining Method to Analyze the Effect of OpenNotes on Clini cal Documentation,

    M. Rahimian, J. L. Warner, S. K. Jain, R. B. Davis, J. A. Zerillo, and R. M. Joyce, “Significant and Distinctive n -Grams in Oncology Notes: A Text-Mining Method to Analyze the Effect of OpenNotes on Clini cal Documentation,” JCO Clin Cancer Inform, vol. 3, pp. 1–9, Jun. 2019,...

  94. [102]

    Scaling up the evaluation of psychotherapy: evaluating motivational interviewing fidelity via statistical text classification,

    D. C. Atkins, M. Steyvers, Z. E. Imel, and P. Smyth, “Scaling up the evaluation of psychotherapy: evaluating motivational interviewing fidelity via statistical text classification,” Implementation Sci, vol. 9, no. 1, p. 49, Dec. 2014, doi: 10.1186/1748-5908-9-49

  95. [103]

    Does Patient Access to Clinical Notes Change Documentation?,

    C. Blease, J. Torous, and M. Hägglund, “Does Patient Access to Clinical Notes Change Documentation?,” Front. Public Health, vol. 8, p. 577896, Nov. 2020, doi: 10.3389/fpubh.2020.577896

  96. [104]

    Presenting time-based risks of stroke and death for Patients facing carotid stenosis treatment options: Patients prefer pie chart s over icon arrays,

    P. Scalia, A. James. O’Malley, M. -A. Durand, P. P. Goodney, and G. Elwyn, “Presenting time-based risks of stroke and death for Patients facing carotid stenosis treatment options: Patients prefer pie chart s over icon arrays,” Patient Education and Counseling , vol. 102, no. 1...

  97. [105]

    How Does Motivational Interviewing Work? Therapist Interpersonal Skill Predicts Client Involvement Within Motivational Interviewing Sessions,

    T. Moyers, W. Miller, and S. Hendrickson, “How Does Motivational Interviewing Work? Therapist Interpersonal Skill Predicts Client Involvement Within Motivational Interviewing Sessions,” Journal of consulting and clinical psychology , vol. 73, p p. 590–8, Aug. 2005, doi: 10.103...

  98. [106]

    The Psychological Meaning of Words: LIWC and Computerized Text Analysis Methods,

    Y. Tausczik and J. Pennebaker, “The Psychological Meaning of Words: LIWC and Computerized Text Analysis Methods,” Journal of Language and Social Psychology , vol. 2 9, pp. 24 –54, Mar. 2010, doi: 10.1177/0261927X09351676

  99. [107]

    Language Style Matching Predicts Relationship Initiation and Stability,

    M. Ireland, R. Slatcher, P. Eastwick, L. Scissors, E. Finkel, and J. Pennebaker, “Language Style Matching Predicts Relationship Initiation and Stability,” Psychological science, vol. 22, pp. 39 –44, Jan. 2011, doi: 10.1177/0956797610392928

  100. [108]

    Multimodal Automatic Coding of Client Behavior in Motivational Interviewing,

    L. Tavabi et al., “Multimodal Automatic Coding of Client Behavior in Motivational Interviewing,” Proceedings of the ... ACM International Conference on Multimodal Interaction. ICMI (Conference), vol. 2020, p. 406, Oct. 2020, doi: 10.1145/3382507.3418853

  101. [109]

    Toward a Theory of Motivational Interviewing,

    W. R. Miller and G. S. Rose, “Toward a Theory of Motivational Interviewing,” Am Psychol, vol. 64, no. 6, pp. 527 –537, Sep. 2009, doi: 10.1037/a0016830

  102. [110]

    A Credit Scoring Model Based on Integrated Mixed Sampling and Ensemble Feature Selection: RBR_XGB,

    X. Lin, Z. Wu, J. Chen, L. Huang, and Z. Shi, “A Credit Scoring Model Based on Integrated Mixed Sampling and Ensemble Feature Selection: RBR_XGB,” Journal of Internet Technology , vol. 23, no. 5, Art. no. 5, Sep. 2022

  103. [111]

    Formality in psychotherapy: How are therapists’ and clients’ use of discourse particles related to therapist empathy?,

    J. H. N. Lee, H. Chui, T. Lee, S. Luk, D. Tao, and N. W. T. Lee, “Formality in psychotherapy: How are therapists’ and clients’ use of discourse particles related to therapist empathy?,” Front Psychiatry, vol. 13, p. 1018170, 2022, doi: 10.3389/fpsyt.2022.1018170

  104. [112]

    Effect of Feature Selection on the Accuracy of Machine Learning Model,

    A. Hamdard and H. Lodin, “Effect of Feature Selection on the Accuracy of Machine Learning Model,” INTERNATIONAL JOURNAL OF MULTIDISCIPLINARY RESEARCH AND ANALYSIS, vol. 06, Sep. 2023, doi: 10.47191/ijmra/v6-i9-66

  105. [113]

    Wu, Interpretable prediction of heart disease based on random forest and SHAP

    L. Wu, Interpretable prediction of heart disease based on random forest and SHAP. 2023. doi: 10.1117/12.2682322

  106. [114]

    Feature selection for machine learning classification problems: A recent overview,

    S. Kotsiantis, “Feature selection for machine learning classification problems: A recent overview,” Artificial Intelligence Review - AIR, vol. 42, Jun. 2011, doi: 10.1007/s10462-011-9230-1

  107. [115]

    A Comparative Analysis of Explainable AI Techniques for Enhanced Model Interpretability,

    S. Y and M. Challa, “A Comparative Analysis of Explainable AI Techniques for Enhanced Model Interpretability,” in 2023 3rd International Conference on Pervasive Computing and Social Networking (ICPCSN), Jun. 2023, pp. 229 –234. doi: 10.1109/ICPCSN58827.2023.00043

  108. [116]

    Symbolic Chain-of-Thought Distillation: Small Models Can Also ‘Think’ Step-by-Step,

    L. H. Li, J. Hessel, Y. Yu, X. Ren, K. -W. Chang, and Y. Choi, “Symbolic Chain-of-Thought Distillation: Small Models Can Also ‘Think’ Step-by-Step,” Ap r. 15, 2024, arXiv: arXiv:2306.14050. doi: 10.48550/arXiv.2306.14050

  109. [117]

    Using ChatGPT as a therapist?,

    Cairo-TenThirteen, “Using ChatGPT as a therapist?,” r/ChatGPTPro. Accessed: Dec. 30, 2024. [Online]. Available: www.reddit.com/r/ChatGPTPro/comments/126rtvb/using_chatgpt_as_a_t herapist/

  110. [118]

    Using motivational interviewing and brief actio n planning for adopting and maintaining positive health behaviors,

    S. A. Cole, D. Sannidhi, Y. T. Jadotte, and A. Rozanski, “Using motivational interviewing and brief actio n planning for adopting and maintaining positive health behaviors,” Progress in Cardiovascular Diseases, vol. 77, pp. 86–94, Mar. 2023, doi: 10.1016/j.pcad.2023.02.003

  111. [119]

    Questions and Reflections: The Use of Motivational Interviewing Microskills in a Peer-Led Brief Alcohol Intervention for College Students,

    S. Tollison, C. Lee, T. Neil, N. Olson, and M. Larimer, “Questions and Reflections: The Use of Motivational Interviewing Microskills in a Peer-Led Brief Alcohol Intervention for College Students,” Behavior therapy, vol. 39, pp. 183–94, Jun. 2008, doi: 10.1016/j.beth.2007.07.001

  112. [120]

    Manual for the Motivational Interviewing Skill Code (MISC),

    P. Amrhein, W. Miller, T. Moyers, and D. Ernst, “Manual for the Motivational Interviewing Skill Code (MISC),” Motivational Interviewing Skill Code v. 2.1 , Jan. 2008, [Online]. Available: https://digitalcommons.montclair.edu/psychology-facpubs/27

  113. [121]

    Analyzing the language of therapist empathy in Motivational Interview based psychotherapy,

    B. Xiao, D. Can, P. G. Georgiou, D. Atkins, and S. S. Narayanan, “Analyzing the language of therapist empathy in Motivational Interview based psychotherapy,” in Proceedings of The 2012 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, Dec...

  114. [122]

    Physician Empathy and Listening: Associations with Patient Satisfaction and Autonomy,

    K. I. Pollak et al., “Physician Empathy and Listening: Associations with Patient Satisfaction and Autonomy,” J Am Board Fam Med, vol. 24, no. 6, pp. 665–672, Nov. 2011, doi: 10.3122/jabfm.2011.06.110025

  115. [123]

    Therapist Empathy Assessment in Motivational Interviews,

    L. Tavabi et al. , “Therapist Empathy Assessment in Motivational Interviews,” in 2023 11th International Conference on Affective Computing and Intelligent Interaction (ACII) , Sep. 2023, pp. 1 –8. doi: 10.1109/ACII59096.2023.10388176

  116. [124]

    MaxMin -RLHF: Alignment with Diverse Human Preferences,

    S. Chakraborty et al. , “MaxMin -RLHF: Alignment with Diverse Human Preferences,” Dec. 26, 2024, arXiv: arXiv:2402.08925. doi: 10.48550/arXiv.2402.08925

  117. [125]

    A Mixed -Method Comparison of Therapist and Client Language across Four Therapeutic Approaches,

    A. H. Qiu and D. Tay, “A Mixed -Method Comparison of Therapist and Client Language across Four Therapeutic Approaches,” Journal of Constructivist Psychology , vol. 36, pp. 1 –24, Jan. 2022, doi: 10.1080/10720537.2021.2021570

  118. [126]

    Language Style Matching, Engagement, and Impasse in Negotiations,

    M. Ireland an d M. Henderson, “Language Style Matching, Engagement, and Impasse in Negotiations,” Negotiation and Conflict Management Research , vol. 7, pp. 1 –16, Feb. 2014, doi: 10.1111/ncmr.12025

  119. [127]

    The alliance,

    A. Horvath, “The alliance,” Psychotherapy Theory Research & Practice, vol. 38, pp. 365 –372, Dec. 2001, doi: 10.1037//0033 - 3204.38.4.365

  120. [128]

    Therapist competence, therapy quality, and therapist training,

    C. Fairburn and Z. Cooper, “Therapist competence, therapy quality, and therapist training,” Behaviour research and therapy, vol. 49, pp. 373– 8, Jun. 2011, doi: 10.1016/j.brat.2011.03.005

  121. [129]

    J. H. N. Lee et al., Durational Patterning at Discourse Boundaries in Relation to Therapist Empathy in Psychotherapy . 2022, p. 5252. doi: 10.21437/Interspeech.2022-722

  122. [130]

    Therapist influence on client language during motivational interviewing sessions,

    T. Moyers and T. Martin, “Therapist influence on client language during motivational interviewing sessions,” Journal of substance abuse treatment, vol. 30, pp. 245–51, May 2006, doi: 10.1016/j.jsat.2005.12.003

  123. [131]

    Barriers to Effective Mental Health Services for African Americans,

    L. R. Snowden, “Barriers to Effective Mental Health Services for African Americans,” Ment Health Serv Res , vol. 3, no. 4, pp. 181 –187, Dec. 2001, doi: 10.1023/A:1013172913880

  124. [132]

    Use of Empathy by Healthcare Professionals Learning Motivational Interviewing: A Qualitative Analysis,

    Krishna Pillai A., “Use of Empathy by Healthcare Professionals Learning Motivational Interviewing: A Qualitative Analysis,” thesis, 2010. Accessed: Dec. 17, 2024. [Online]. Available: https://etd.auburn.edu//handle/10415/2112

  125. [133]

    Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback,

    Y. Bai et al. , “Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback,” Apr. 12, 2022, arXiv: arXiv:2204.05862. doi: 10.48550/arXiv.2204.05862

  126. [134]

    Training language models to follow instructions with human feedb ack,

    L. Ouyang et al., “Training language models to follow instructions with human feedb ack,” Mar. 04, 2022, arXiv: arXiv:2203.02155. doi: 10.48550/arXiv.2203.02155

  127. [135]

    M. J. Lambert, S. L. Garfield, and A. E. Bergin, Eds., Bergin and Garfield’s handbook of psychotherapy and behavior change, Sixth edition. Hoboken, N.J: Wiley, 2013

  128. [136]

    A Roadmap to Pluralistic Alignment,

    T. Sorensen et al., “A Roadmap to Pluralistic Alignment,” Aug. 20, 2024, arXiv: arXiv:2402.05070. doi: 10.48550/arXiv.2402.05070

  129. [137]

    Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative AI,

    M. Abbasian et al., “Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative AI,” npj Digit. Med., vol. 7, no. 1, pp. 1–14, Mar. 2024, doi: 10.1038/s41746-024-01074-z

  130. [138]

    Towards Understanding Sycophancy in Language Models,

    M. Sharma et al., “Towards Understanding Sycophancy in Language Models,” Oct. 27, 2023, arXiv: arXiv:2310.13548. doi: 10.48550/arXiv.2310.13548

  131. [139]

    Challenges of Large Language Models for Mental Health Counseling,

    N. C. Chung, G. Dyer, and L. Brocki, “Challenges of Large Language Models for Mental Health Counseling,” Nov. 23, 2023, arXiv: arXiv:2311.13857. doi: 10.48550/arXiv.2311.13857

  132. [140]

    The dissemination of empirically supported treatments: A view to the future,

    D. H. Barlow, J. T. Levitt, and L. F. Bufka, “The dissemination of empirically supported treatments: A view to the future,” Behaviour Research and Therapy , vol. 37, no. Suppl 1, pp. S147 –S162, 1999, doi: 10.1016/S0005-7967(99)00054-6

  133. [141]

    Testing the integrity of a psychotherapy protocol: Assessment of adherence and competence,

    J. Waltz, M. E. Addis, K. Koerner, and N. S. Jacobson, “Testing the integrity of a psychotherapy protocol: Assessment of adherence and competence,” Journal of Consulting and Clinical Psychology, vol. 61, no. 4, pp. 620–630, 1993, doi: 10.1037/0022-006X.61.4.620

  134. [142]

    The effectiveness and applicability of motivational interviewing: a practice‐friendly review of four meta‐ analyses,

    B. Lundahl and B. L. Burke, “The effectiveness and applicability of motivational interviewing: a practice‐friendly review of four meta‐ analyses,” J Clin Psychol, vol. 65, no. 11, pp. 1232–1245, Nov. 2009, doi: 10.1002/jclp.20638

  135. [143]

    Culture and motivational interviewing,

    H. Oh and C. Lee, “Culture and motivational interviewing,” Patient Educ Couns , vol. 99, no. 11, pp. 1914 –1919, Nov. 2016, doi: 10.1016/j.pec.2016.06.010

  136. [145]

    Mental Health: Culture, Race, and Ethnicity —A Supplement to Mental Health: A Report of the Surgeon General,

    D. Satcher, “Mental Health: Culture, Race, and Ethnicity —A Supplement to Mental Health: A Report of the Surgeon General,” 2001, Accessed: Dec. 12, 2024. [Online]. Available: http://hdl.handle.net/1903/22834

  137. [146]

    Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning,

    Z. Xie et al., “Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning,” Mar. 23, 20 24, arXiv: arXiv:2403.15737. doi: 10.48550/arXiv.2403.15737

  138. [147]

    ELIZA —a computer program for the study of natural language communication between man and machine,

    J. Weizenbaum, “ELIZA —a computer program for the study of natural language communication between man and machine,” Commun. ACM, vol. 9, no. 1, pp. 36–45, 1966, doi: 10.1145/365153.365168

  139. [148]

    The ethics of ChatGPT in medicine and healthcare: a systematic review on Large Language Models (LLMs),

    J. Haltaufderheide and R. Ranisch, “The ethics of ChatGPT in medicine and healthcare: a systematic review on Large Language Models (LLMs),” npj Digit. Med. , vol. 7, no. 1, pp. 1 –11, Jul. 2024, doi: 10.1038/s41746-024-01157-x

  140. [149]

    ChatGPT and global public health: Applications, challenges, ethical considerations and mitigation strategies,

    A. A. Parray, Z. M. Inam, D. Ramonfaur, S. S. Haider, S. K. Mistry, and A. K. Pandya, “ChatGPT and global public health: Applications, challenges, ethical considerations and mitigation strategies,” Global Transitions, vol. 5, pp. 50–54, 2023, doi: 10.1016/j.glt.2023.05.001

  141. [151]

    A framework for language technologies in behavioral research and clinical applications: Ethical challenges, implications, and solutions.,

    C. Diaz-Asper, M. K. Hauglid, C. Chandler, A. S. Cohen, P. W. Foltz, and B. Elvevåg, “A framework for language technologies in behavioral research and clinical applications: Ethical challenges, implications, and solutions.,” American Psychologist, vol. 79, no. 1, pp. 79 –91, J...

  142. [152]

    Taxonomy of Risks posed by Language Models,

    L. Weidinger et al., “Taxonomy of Risks posed by Language Models,” in 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul Republic of Korea: ACM, Jun. 2022, pp. 214 –229. doi: 10.1145/3531146.3533088

  143. [2024]

    Available: https://aclanthology.org/E06-1032

    [Online]. Available: https://aclanthology.org/E06-1032

  144. [2252]

    doi: 10.18653/v1/D17-1238

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.