Pith. sign in

REVIEW 3 cited by

Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.14385 v4 pith:PB7KH2DR submitted 2023-07-26 cs.CL

classification cs.CL
keywords healthllmsmentaltasksmodelscapabilitygpt-4language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advances in large language models (LLMs) have empowered a variety of applications. However, there is still a significant gap in research when it comes to understanding and enhancing the capabilities of LLMs in the field of mental health. In this work, we present a comprehensive evaluation of multiple LLMs on various mental health prediction tasks via online text data, including Alpaca, Alpaca-LoRA, FLAN-T5, GPT-3.5, and GPT-4. We conduct a broad range of experiments, covering zero-shot prompting, few-shot prompting, and instruction fine-tuning. The results indicate a promising yet limited performance of LLMs with zero-shot and few-shot prompt designs for mental health tasks. More importantly, our experiments show that instruction finetuning can significantly boost the performance of LLMs for all tasks simultaneously. Our best-finetuned models, Mental-Alpaca and Mental-FLAN-T5, outperform the best prompt design of GPT-3.5 (25 and 15 times bigger) by 10.9% on balanced accuracy and the best of GPT-4 (250 and 150 times bigger) by 4.8%. They further perform on par with the state-of-the-art task-specific language model. We also conduct an exploratory case study on LLMs' capability on mental health reasoning tasks, illustrating the promising capability of certain models such as GPT-4. We summarize our findings into a set of action guidelines for potential methods to enhance LLMs' capability for mental health tasks. Meanwhile, we also emphasize the important limitations before achieving deployability in real-world mental health settings, such as known racial and gender bias. We highlight the important ethical risks accompanying this line of research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Gold Standard Dataset and Evaluation Framework for Depression Detection and Explanation in Social Media using LLMs

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new expert-annotated dataset with span-level depression labels is used to compare GPT-4.1, Claude 3.7, and Gemini 2.5 Pro on explanation faithfulness, finding no consistent gain from few-shot prompting.

  2. Systematic Evaluation of Machine-Generated Reasoning and PHQ-9 Labeling for Depression Detection Using Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A subtask decomposition of depression detection shows LLMs are biased by explicit depression keywords, and DPO fine-tuning on quality-filtered machine-generated rationales improves joint PHQ-9 labeling on the hardest samples.

  3. SELF-PERCEPT: Introspection Improves Large Language Models' Detection of Multi-Person Mental Manipulation in Conversations

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A two-stage introspection prompt, SELF-PERCEPT, modestly improves LLM detection of mental manipulation in multi-party dialogues on a new 220-dialogue dataset.

Pith tools