Pith. sign in

REVIEW 2 cited by

A Gold Standard Dataset and Evaluation Framework for Depression Detection and Explanation in Social Media using LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.19899 v1 pith:AUXFSRED submitted 2025-07-26 cs.CL

A Gold Standard Dataset and Evaluation Framework for Depression Detection and Explanation in Social Media using LLMs

classification cs.CL
keywords datasetdepressionevaluationllmsmediasocialdetectionexplanation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Early detection of depression from online social media posts holds promise for providing timely mental health interventions. In this work, we present a high-quality, expert-annotated dataset of 1,017 social media posts labeled with depressive spans and mapped to 12 depression symptom categories. Unlike prior datasets that primarily offer coarse post-level labels \cite{cohan-etal-2018-smhd}, our dataset enables fine-grained evaluation of both model predictions and generated explanations. We develop an evaluation framework that leverages this clinically grounded dataset to assess the faithfulness and quality of natural language explanations generated by large language models (LLMs). Through carefully designed prompting strategies, including zero-shot and few-shot approaches with domain-adapted examples, we evaluate state-of-the-art proprietary LLMs including GPT-4.1, Gemini 2.5 Pro, and Claude 3.7 Sonnet. Our comprehensive empirical analysis reveals significant differences in how these models perform on clinical explanation tasks, with zero-shot and few-shot prompting. Our findings underscore the value of human expertise in guiding LLM behavior and offer a step toward safer, more transparent AI systems for psychological well-being.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. WPG-MoE: Weak-Prior-Guided Dense Mixture-of-Experts for User-Level Social Media Depression Detection

    cs.CL 2026-07 conditional novelty 6.0

    A dense mixture-of-experts model guided by training-only weak evidence-layout priors outperforms single-detector baselines on Chinese and English user-level depression detection.

  2. Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

    cs.AI 2026-07 conditional novelty 5.0

    An LLM-assisted, expert-verified annotation pipeline produces DSM-5-TR-aligned depression labels with evidence traces, showing high agreement and reduced effort in a 10-case pilot, while its self-evolving memory remai...