REVIEW 4 major objections 6 minor 23 references
What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma
T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims a new expert-annotated interview corpus, built on attribution theory, lets researchers detect fine-grained mental-health stigma in conversation and shows that state-of-the-art language models still miss its most subtle…
desk verdict A genuinely new and useful interview-based stigma corpus, but the theory-grounded label claim needs construct validation and the mixed-response annotation rule likely over-attributes stigma. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the MHS TIGMA INTERVIEW corpus itself, a collection of 4,141 expert-annotated snippets from chatbot interviews, and the attribution model that shapes it. The annotation scheme condenses the attribution model and AQ-27 into seven stigma attributions, each linked to one interview question embedded in a vignette about the fictional character Avery. The interview protocol's role is to elicit candid reactions with scenario-embedded questions, randomized order, follow-up prompts, and neutral self-disclosure from the chatbot, and the benchmark experiments then use the corpus to test how well models can recover the human labels.
What would settle it
Administer the original AQ-27 to the same 684 participants shortly after their chatbot interview and compute person-level correlation between survey scores and the stigma labels assigned to their interview snippets; if the correlation is negligible, the interview adaptation is not measuring the same construct. A complementary check would compare chatbot-elicited responses with anonymous written self-reports on the same vignette to detect social-desirability suppression.
Extended reading notes
Core claim
The central claim is that mental-health stigma can be operationalized into seven measurable attribution types and that a carefully elicited interview corpus can surface real stigmatizing talk that existing datasets and models miss. The authors adapt the Attribution Questionnaire-27 into seven scenario-based chatbot questions about a fictional person with depression, collect responses from 684 screened English-speaking adults, and have two expert annotators label each snippet under a codebook informed by the attribution model, reaching 0.71 Cohen's kappa. The resulting corpus labels 46% of snippets as stigmatizing, with responsibility and social distance the most common types, and with toxicity lower than in hate-speech benchmarks, evidence that stigma here appears as veiled, normalized, or well-intentioned language. Benchmarks on eight-way classification show that larger instruction-tuned models improve with a full codebook but still fall short, with GPT-4o's best macro F1 around 0.757, and a qualitative analysis of GPT-4o's 137 misclassified snippets identifies distancing language, misuse of psychiatric terms, coercive phrasing, differential support, patronization, and minimization as the recurring failure modes.
Load-bearing premise
The load-bearing premise is that adapting the Attribution Questionnaire-27 into seven chatbot interview questions about a fictional person measures the same mental-health stigma construct as the original survey, and that participants answered candidly rather than editing their responses to appear socially acceptable.
Editorial extensions
If this is right
- Systems can be trained to classify stigma by its underlying attribution (blame, distance, anger, withholding help, pity, forced treatment, fear) instead of treating it as a single toxic category.
- Researchers get a benchmark where the full codebook given to human annotators measurably improves LLM performance, so gains from better instructions can be tracked against expert labels.
- The documented demographic and geographic metadata make it possible to study how stigma expression varies with gender, country, education, and personal exposure to mental illness.
- The corpus can be used to have models role-play interviewees and compare their generated responses with the human ones, exposing whether models internalize and perpetuate the same stigmatizing attributions.
Reading between the lines
- Testable extension: give the same participants the original AQ-27 survey and correlate their scores with the stigma labels from their interview snippets; a weak correlation would suggest the chatbot adaptation changes the construct being measured.
- Because each interview snippet is tied to the attribution-specific question that elicited it, prompt-ablating that question text would test whether model errors come from the wording of the question rather than from the participant's stigma.
- The same interview and annotation protocol could be run in non-Western languages and cultures; the current corpus, sampled mostly from Western English-speaking participants, cannot show whether the seven attribution types structure stigma expression everywhere.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces MHS TIGMA INTERVIEW, an expert-annotated, theory-informed corpus for mental-health stigma detection, comprising 4,141 interview snippets from 684 participants who interacted with a chatbot about a vignette character with depression. Seven interview questions adapt items from the Attribution Questionnaire-27 (AQ-27) to probe responsibility, social distance, anger, helping, pity, coercive segregation, and fear. Two trained annotators labeled each snippet with one of seven stigma attributions or a non-stigmatizing category, with checkpoint kappas and a final Cohen's kappa of 0.71. The authors benchmark fine-tuned RoBERTa and several large language models under zero-shot, one-shot, and full-codebook prompting, and qualitatively analyze 137 GPT-4o misclassifications to characterize subtle stigmatizing language. The paper claims to provide the first large-scale, open-source mental-health stigma interview dataset with theoretical grounding and documented socio-cultural backgrounds, and concludes that neural models alone remain insufficient for stigma detection.
Significance. If the corpus labels are valid, this is a genuinely useful resource. It addresses a real gap left by social-media and synthetic corpora, uses a recognized psychological framework, and ships with unusually thorough documentation: IRB approval, consent procedures, iterative codebook development, checkpoint agreement, an agreement matrix, annotator feedback, and access-controlled release. The benchmark experiments and the qualitative error analysis are valuable starting points for future work. The main risk is construct validity: the adapted interview protocol and the annotation rules may over-attribute stigma, and this would propagate into every benchmark number and qualitative conclusion. The paper deserves serious consideration, but the labeling assumptions need to be tested or explicitly softened before the strongest claims can be accepted.
major comments (4)
- [3.1, Table 2, Limitations] The operationalization of the AQ-27 as seven open-ended chatbot questions is supported only by face validity; no concurrent or discriminant validation against the original AQ-27 or another stigma scale is reported. This is load-bearing because the corpus's claim to be 'theory-grounded' and the benchmark conclusions in Section 4.2 both rest on the assumption that the chatbot questions elicit the same construct as the AQ-27. I would ask for a small validation subsample (e.g., administering the AQ-27 to a subset of participants) or, at minimum, an explicit reframing of the claim to 'theory-informed' and a discussion of the validation gap. The Limitations section does not currently acknowledge this missing validation.
- [Appendix H.3, Figure 1] Annotation rule 1 in the codebook ('If a message mentions both stigmatized and non-stigmatized statements, annotate it as Stigmatizing') systematically resolves mixed conversational responses toward stigma. Figure 1 illustrates the risk: the participant ultimately says they would rent to Avery, yet the snippet is labeled Stigmatizing (Social Distance) because of the preceding hesitation. Because these labels are the ground truth for Table 4 and the qualitative claim in Section 4.3, this rule can overstate stigma prevalence and inflate apparent model failures. I recommend either refining the rule for hedged responses or reporting sensitivity analyses with a more conservative rule.
- [3.3, Limitations] The final labels rest on two annotators with similar demographic backgrounds (both Asian and in their twenties), with a moderate Cohen's kappa of 0.71, and disagreements are resolved by consensus. The Limitations section candidly acknowledges annotator subjectivity and promises annotator-level labels only in a future v2.0, but the released corpus and all benchmark results use the consensus labels. For a construct as culturally and socially variable as stigma, I would like to see at least a disagreement analysis or an external audit by additional annotators from different backgrounds before treating these labels as reliable ground truth.
- [3.2.1, 3.2.2] The interview format and participant screening introduce assumptions that are not validated. Excluding participants with immediate mental-health concerns via the K6, and asking about willingness to rent, help, or hospitalize a fictional character through a chatbot, may suppress or redirect the very attitudes the corpus aims to measure. Section 3.2.1 cites scenario-embedded questions and randomized order to mitigate social-desirability bias, but no evidence is provided that the chatbot setting elicits candid stigmatizing attitudes. This is connected to the construct-validity concern above and should be addressed directly, for example by comparing chatbot interview responses with a standard survey administration in a pilot study.
minor comments (6)
- [Abstract, Section 3.4] Demographic documentation is based on 555 of 684 participants (81.1%); the abstract's phrase 'documented socio-cultural backgrounds' should be qualified to avoid overstating coverage.
- [Figure 1] The snippet in Figure 1 appears with three labels (Stigmatizing (Responsibility), Stigmatizing (Social Distance), Non-stigmatizing), while Section 3.3 describes a single-label task; please clarify whether this is a multi-label illustration or an error in the figure.
- [Table 1] The entry for Roesler et al. uses '1' in the Theory-Grounded column with no footnote; use checkmarks or explain the partial mark.
- [Equation (1), Section 3.2.1] The 25- and 150-character thresholds in the follow-up question policy are justified only by an 8-participant pilot; a sentence on the robustness of these thresholds would be helpful.
- [Appendix H.3] The keyword 'distance' appears in both Social Distance and Coercive Segregation definitions, which may confuse both annotators and prompted models; consider disambiguating the two sets of keywords.
- [Section 4.1] GPT-4o was evaluated with a single run and a single temperature (0.2) due to budget constraints; the reported F1 differences should be interpreted with this uncertainty in mind, and a caveat in the text would help.
Circularity Check
No significant circularity: the corpus, annotations, and benchmarks are self-contained against external theory and human labels.
full rationale
This is a dataset-construction and benchmarking paper rather than a derivation. The labels are produced by two expert annotators guided by an external theoretical framework (Corrigan et al., 2003) via the AQ-27 operationalization in Table 2; no label or benchmark number is defined in terms of the paper's own outputs or conclusions. The benchmark compares model predictions against the human-annotated labels, which is standard supervised evaluation, and the qualitative error analysis in Section 4.3 is a post hoc reading of GPT-4o's 137 misclassifications rather than a claim derived from the annotation scheme. The author self-citations (Lee et al., 2023; Meng et al., 2024) are used to justify combining AQ-27 subscales and interview structure; those are design decisions, not the paper's predicted result, and they do not force any finding. The central claims are factual corpus properties (4,141 snippets from 684 participants, Cohen's kappa of 0.71) and model performance numbers, all verifiable from the released data and standard evaluation protocols. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no theoretical result is asserted by definition. Accordingly, no circular step is present.
Assumptions & free parameters
free parameters (1)
- Follow-up question length thresholds =
25 and 150 characters
assumptions (5)
- domain assumption Attribution theory (Kelley, 1967) and Corrigan et al.'s (2003) attribution model provide a valid decomposition of public mental-health stigma.
- domain assumption The AQ-27 (Corrigan, 2012), originally a survey instrument, can be validly adapted to an open-ended chatbot interview without changing the measured construct.
- domain assumption Chatbot-facilitated interviews with a fictional vignette elicit genuine stigmatizing attitudes rather than only socially desirable responses.
- domain assumption Expert-guided annotation by two research assistants, with consensus discussion and a Cohen's kappa of 0.71, produces reliable ground-truth labels for the benchmark.
- ad hoc to paper Excluding participants with immediate mental-health concerns (K6 screening) does not materially distort the stigma distribution the corpus intends to represent.
Cite this review
Pith. "Pith review of What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma." pith.science (2026). https://pith.science/paper/M7LDYUW7
@misc{pith2026250512727,
author = {Pith},
title = {Pith review of: What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma},
year = {2026},
howpublished = {\url{https://pith.science/paper/M7LDYUW7}},
note = {Machine review of arXiv:2505.12727}
}
read the original abstract
Mental-health stigma remains a pervasive social problem that hampers treatment-seeking and recovery. Existing resources for training neural models to finely classify such stigma are limited, relying primarily on social-media or synthetic data without theoretical underpinnings. To remedy this gap, we present an expert-annotated, theory-informed corpus of human-chatbot interviews, comprising 4,141 snippets from 684 participants with documented socio-cultural backgrounds. Our experiments benchmark state-of-the-art neural models and empirically unpack the challenges of stigma detection. This dataset can facilitate research on computationally detecting, neutralizing, and counteracting mental-health stigma. Our corpus is openly available at https://github.com/HanMeng2004/Mental-Health-Stigma-Interview-Corpus.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
36 Keywords: responsible, responsibility, weakness, stem from, blame, etc
Responsibility: A common stigmatized thought is that people have control over and are responsible for their mental illness and related symptoms. 36 Keywords: responsible, responsibility, weakness, stem from, blame, etc
-
[2]
Keywords: worried, apprehensive, upset, unhappy, worrying, get along, distance, etc
Social Distance: The constant stigmatized behavior is to stay away from people with mental illness. Keywords: worried, apprehensive, upset, unhappy, worrying, get along, distance, etc
-
[3]
Keywords: annoyed, angry, anger, irritated, lost temper, etc
Anger: The stigmatized thought is to be irritated or annoyed because people are blamed for their mental illness. Keywords: annoyed, angry, anger, irritated, lost temper, etc
-
[4]
Keywords: lack, hinder, refuse, less inclined to, help, assist, etc
Helping: The stigmatized behavior is withholding support towards them because of biased thoughts. Keywords: lack, hinder, refuse, less inclined to, help, assist, etc
-
[5]
Keywords: concern, pity, sympathy, tough, sad, lack, etc
Pity: A common stigmatized thought is to be unsympathetic towards people with mental illness, and place blame on them for their mental health challenges. Keywords: concern, pity, sympathy, tough, sad, lack, etc
-
[6]
Coercive Segregation: The stigmatized behavior is to send people with mental illness to institutions away from their community and force people with mental illness to participate in medication management or other treatments. Keywords: separate, lonely, warrant, hospitalization, distance, treatment, respect, neighborhood, neighbor, socialization, companion...
-
[7]
Fear: The stigmatized thought is to believe that people with mental illness are not safe or feel frightened since people with mental illness are dangerous or unpredictable. Keywords: violence, threatened, safe, danger, afraid, cautious, hurt, safety, frightened, scare, intimidate, threatening, anger, fear, yelling, etc
-
[8]
Please carefully read the additional rules and adhere strictly to them when annotating the data:
Non-stigmatizing: The text explicitly conveys non-stigmatizing content, displaying positive attitudes such as sympathy, assistance, friendliness, etc. Please carefully read the additional rules and adhere strictly to them when annotating the data:
Show all 23 references
-
[9]
If it doesn't have any stigmatized thoughts, choose H. ''' 38
-
[15]
If a message mentions both stigmatized and non-stigmatized statements, annotate it as Stigmatizing
-
[16]
People have control over and are responsible for their mental illness and related symptoms
When discussing ‘responsibility,’ common stigmatized thoughts are: i. People have control over and are responsible for their mental illness and related symptoms. ii. They think their situation is a result of their actions. iii. They think it’s a personal weakness. iv. They can...
-
[17]
Preferring to stay away from people with mental illness
When discussing ‘social distance,’ common stigmatized thoughts are: vi. Preferring to stay away from people with mental illness. vii. Associating mental illness with negative living situations or predicting future conflicts. 37 viii. Rejecting the idea of living with people wi...
-
[18]
Feeling irritated or annoyed with individuals because of their mental illness
When discussing 'anger', common stigmatized thoughts are: x. Feeling irritated or annoyed with individuals because of their mental illness. xi. Feeling uncomfortable around people with mental illness. xii. Expressing ‘dislike’ for people with mental illness. xiii. Other statem...
-
[19]
Withholding support or assistance towards people with mental illness
When discussing 'helping', common stigmatized thoughts are: xiv. Withholding support or assistance towards people with mental illness. xv. Offering indirect help, e.g., guidance, to people with mental illness. xvi. Offering conditional help to people with mental illness. xvii....
-
[20]
Feeling no concern or sympathy for individuals with mental illness
When discussing 'pity', common stigmatized thoughts are: xix. Feeling no concern or sympathy for individuals with mental illness. xx. Other statements reflect the mean opinions. If a message contains the stigmatized thought mentioned in xix. - xx., choose E
-
[21]
Suggesting for sending people with mental illness to institutions away from their community
When discussing 'coercive segregation', common stigmatized thoughts are: xxi. Suggesting for sending people with mental illness to institutions away from their community. xxii. Suggesting that people with mental illness need to take the medication or treatment, with or without...
-
[22]
Feeling frightened of people with mental illness
When discussing 'fear', common stigmatized thoughts are: xxvi. Feeling frightened of people with mental illness. xxvii. Viewing people with mental illness as dangerous and unpredictable. xxviii. feeling extra cautious when interacting with people with mental illness. xxix. Ass...
-
[1964]
how" or “what
– that language might be able to capture. B More Details about Data Collection: Chatbot-based Interview B.1 Vignettes The clinical version appears below: Avery is employed by a company, and in their spare time, they are dedicated to lifelong learning, doing extensive reading a...
2024
-
[2020]
i hear you, i feel you
Do models of mental health based on social media data generalize? In Findings of the Associa- tion for Computational Linguistics: EMNLP 2020 , pages 3774–3788, Online. Association for Computa- tional Linguistics. Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, D...
2020 arXiv
-
[2021]
HateCheck: Functional tests for hate speech detection models. In Proceedings of the 59th An- nual Meeting of the Association for Computational Linguistics and the 11th International Joint Confer- ence on Natural Language Processing (Volume 1: Long Papers), pages 41–58, Online....
2014
-
[2022]
In Proceedings of the 2022 Con- ference on Empirical Methods in Natural Language Processing, pages 2152–2170, Abu Dhabi, United Arab Emirates
Gendered mental health stigma in masked language models. In Proceedings of the 2022 Con- ference on Empirical Methods in Natural Language Processing, pages 2152–2170, Abu Dhabi, United Arab Emirates. Association for Computational Lin- guistics. Bruce G. Link, Francis T. Cullen...
2022 arXiv
-
[2023]
In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 14523–14530
Everyone’s voice matters: Quantifying anno- tation disagreement using demographic information. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 14523–14530. Zeerak Waseem and Dirk Hovy. 2016. Hateful symbols or hateful people? predictive featu...
2016 arXiv
-
[2024]
social prim- ing
Exploring the relationship between intrin- sic stigma in masked language models and train- ing data using the stereotype content model. In Proceedings of the Fifth Workshop on Resources and ProcessIng of linguistic, para-linguistic and extra-linguistic Data from people with va...
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.