Pith. sign in

REVIEW 4 major objections 6 minor 23 references

What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma

T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims a new expert-annotated interview corpus, built on attribution theory, lets researchers detect fine-grained mental-health stigma in conversation and shows that state-of-the-art language models still miss its most subtle…

desk verdict A genuinely new and useful interview-based stigma corpus, but the theory-grounded label claim needs construct validation and the mixed-response annotation rule likely over-attributes stigma. read the letter →

arxiv 2505.12727 v2 pith:M7LDYUW7 submitted 2025-05-19 cs.CL cs.CYcs.HC

classification cs.CLcs.CYcs.HC
keywords mental-healthstigmainterviewcorpusattributiontheorychatbotmulti-labelannotationdetectionbenchmarkexpertlargelanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to give computational research on mental-health stigma a data source it has lacked: a large, open, theory-grounded corpus of how ordinary people actually talk about stigma in conversation. It introduces 4,141 expert-annotated transcript snippets drawn from 684 participants in structured human-chatbot interviews, each snippet labeled as non-stigmatizing or as one of seven attribution types (responsibility, social distance, anger, helping, pity, coercive segregation, fear) derived from a standard attribution model of public stigma. The authors argue that the corpus captures veiled, socially acceptable forms of stigma that social-media and synthetic corpora miss, and they benchmark open models and GPT-4o to show that state-of-the-art neural classifiers still misread these subtleties. If the corpus does what it claims, it supplies a shared benchmark for detecting, neutralizing, and counteracting mental-health stigma in language.

What carries the argument

The load-bearing object is the MHS TIGMA INTERVIEW corpus itself, a collection of 4,141 expert-annotated snippets from chatbot interviews, and the attribution model that shapes it. The annotation scheme condenses the attribution model and AQ-27 into seven stigma attributions, each linked to one interview question embedded in a vignette about the fictional character Avery. The interview protocol's role is to elicit candid reactions with scenario-embedded questions, randomized order, follow-up prompts, and neutral self-disclosure from the chatbot, and the benchmark experiments then use the corpus to test how well models can recover the human labels.

What would settle it

Administer the original AQ-27 to the same 684 participants shortly after their chatbot interview and compute person-level correlation between survey scores and the stigma labels assigned to their interview snippets; if the correlation is negligible, the interview adaptation is not measuring the same construct. A complementary check would compare chatbot-elicited responses with anonymous written self-reports on the same vignette to detect social-desirability suppression.

Watch

Extended reading notes

Core claim

The central claim is that mental-health stigma can be operationalized into seven measurable attribution types and that a carefully elicited interview corpus can surface real stigmatizing talk that existing datasets and models miss. The authors adapt the Attribution Questionnaire-27 into seven scenario-based chatbot questions about a fictional person with depression, collect responses from 684 screened English-speaking adults, and have two expert annotators label each snippet under a codebook informed by the attribution model, reaching 0.71 Cohen's kappa. The resulting corpus labels 46% of snippets as stigmatizing, with responsibility and social distance the most common types, and with toxicity lower than in hate-speech benchmarks, evidence that stigma here appears as veiled, normalized, or well-intentioned language. Benchmarks on eight-way classification show that larger instruction-tuned models improve with a full codebook but still fall short, with GPT-4o's best macro F1 around 0.757, and a qualitative analysis of GPT-4o's 137 misclassified snippets identifies distancing language, misuse of psychiatric terms, coercive phrasing, differential support, patronization, and minimization as the recurring failure modes.

Load-bearing premise

The load-bearing premise is that adapting the Attribution Questionnaire-27 into seven chatbot interview questions about a fictional person measures the same mental-health stigma construct as the original survey, and that participants answered candidly rather than editing their responses to appear socially acceptable.

Editorial extensions

If this is right

  • Systems can be trained to classify stigma by its underlying attribution (blame, distance, anger, withholding help, pity, forced treatment, fear) instead of treating it as a single toxic category.
  • Researchers get a benchmark where the full codebook given to human annotators measurably improves LLM performance, so gains from better instructions can be tracked against expert labels.
  • The documented demographic and geographic metadata make it possible to study how stigma expression varies with gender, country, education, and personal exposure to mental illness.
  • The corpus can be used to have models role-play interviewees and compare their generated responses with the human ones, exposing whether models internalize and perpetuate the same stigmatizing attributions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Testable extension: give the same participants the original AQ-27 survey and correlate their scores with the stigma labels from their interview snippets; a weak correlation would suggest the chatbot adaptation changes the construct being measured.
  • Because each interview snippet is tied to the attribution-specific question that elicited it, prompt-ablating that question text would test whether model errors come from the wording of the question rather than from the participant's stigma.
  • The same interview and annotation protocol could be run in non-Western languages and cultures; the current corpus, sampled mostly from Western English-speaking participants, cannot show whether the seven attribution types structure stigma expression everywhere.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript introduces MHS TIGMA INTERVIEW, an expert-annotated, theory-informed corpus for mental-health stigma detection, comprising 4,141 interview snippets from 684 participants who interacted with a chatbot about a vignette character with depression. Seven interview questions adapt items from the Attribution Questionnaire-27 (AQ-27) to probe responsibility, social distance, anger, helping, pity, coercive segregation, and fear. Two trained annotators labeled each snippet with one of seven stigma attributions or a non-stigmatizing category, with checkpoint kappas and a final Cohen's kappa of 0.71. The authors benchmark fine-tuned RoBERTa and several large language models under zero-shot, one-shot, and full-codebook prompting, and qualitatively analyze 137 GPT-4o misclassifications to characterize subtle stigmatizing language. The paper claims to provide the first large-scale, open-source mental-health stigma interview dataset with theoretical grounding and documented socio-cultural backgrounds, and concludes that neural models alone remain insufficient for stigma detection.

Significance. If the corpus labels are valid, this is a genuinely useful resource. It addresses a real gap left by social-media and synthetic corpora, uses a recognized psychological framework, and ships with unusually thorough documentation: IRB approval, consent procedures, iterative codebook development, checkpoint agreement, an agreement matrix, annotator feedback, and access-controlled release. The benchmark experiments and the qualitative error analysis are valuable starting points for future work. The main risk is construct validity: the adapted interview protocol and the annotation rules may over-attribute stigma, and this would propagate into every benchmark number and qualitative conclusion. The paper deserves serious consideration, but the labeling assumptions need to be tested or explicitly softened before the strongest claims can be accepted.

major comments (4)
  1. [3.1, Table 2, Limitations] The operationalization of the AQ-27 as seven open-ended chatbot questions is supported only by face validity; no concurrent or discriminant validation against the original AQ-27 or another stigma scale is reported. This is load-bearing because the corpus's claim to be 'theory-grounded' and the benchmark conclusions in Section 4.2 both rest on the assumption that the chatbot questions elicit the same construct as the AQ-27. I would ask for a small validation subsample (e.g., administering the AQ-27 to a subset of participants) or, at minimum, an explicit reframing of the claim to 'theory-informed' and a discussion of the validation gap. The Limitations section does not currently acknowledge this missing validation.
  2. [Appendix H.3, Figure 1] Annotation rule 1 in the codebook ('If a message mentions both stigmatized and non-stigmatized statements, annotate it as Stigmatizing') systematically resolves mixed conversational responses toward stigma. Figure 1 illustrates the risk: the participant ultimately says they would rent to Avery, yet the snippet is labeled Stigmatizing (Social Distance) because of the preceding hesitation. Because these labels are the ground truth for Table 4 and the qualitative claim in Section 4.3, this rule can overstate stigma prevalence and inflate apparent model failures. I recommend either refining the rule for hedged responses or reporting sensitivity analyses with a more conservative rule.
  3. [3.3, Limitations] The final labels rest on two annotators with similar demographic backgrounds (both Asian and in their twenties), with a moderate Cohen's kappa of 0.71, and disagreements are resolved by consensus. The Limitations section candidly acknowledges annotator subjectivity and promises annotator-level labels only in a future v2.0, but the released corpus and all benchmark results use the consensus labels. For a construct as culturally and socially variable as stigma, I would like to see at least a disagreement analysis or an external audit by additional annotators from different backgrounds before treating these labels as reliable ground truth.
  4. [3.2.1, 3.2.2] The interview format and participant screening introduce assumptions that are not validated. Excluding participants with immediate mental-health concerns via the K6, and asking about willingness to rent, help, or hospitalize a fictional character through a chatbot, may suppress or redirect the very attitudes the corpus aims to measure. Section 3.2.1 cites scenario-embedded questions and randomized order to mitigate social-desirability bias, but no evidence is provided that the chatbot setting elicits candid stigmatizing attitudes. This is connected to the construct-validity concern above and should be addressed directly, for example by comparing chatbot interview responses with a standard survey administration in a pilot study.
minor comments (6)
  1. [Abstract, Section 3.4] Demographic documentation is based on 555 of 684 participants (81.1%); the abstract's phrase 'documented socio-cultural backgrounds' should be qualified to avoid overstating coverage.
  2. [Figure 1] The snippet in Figure 1 appears with three labels (Stigmatizing (Responsibility), Stigmatizing (Social Distance), Non-stigmatizing), while Section 3.3 describes a single-label task; please clarify whether this is a multi-label illustration or an error in the figure.
  3. [Table 1] The entry for Roesler et al. uses '1' in the Theory-Grounded column with no footnote; use checkmarks or explain the partial mark.
  4. [Equation (1), Section 3.2.1] The 25- and 150-character thresholds in the follow-up question policy are justified only by an 8-participant pilot; a sentence on the robustness of these thresholds would be helpful.
  5. [Appendix H.3] The keyword 'distance' appears in both Social Distance and Coercive Segregation definitions, which may confuse both annotators and prompted models; consider disambiguating the two sets of keywords.
  6. [Section 4.1] GPT-4o was evaluated with a single run and a single temperature (0.2) due to budget constraints; the reported F1 differences should be interpreted with this uncertainty in mind, and a caveat in the text would help.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the corpus, annotations, and benchmarks are self-contained against external theory and human labels.

full rationale

This is a dataset-construction and benchmarking paper rather than a derivation. The labels are produced by two expert annotators guided by an external theoretical framework (Corrigan et al., 2003) via the AQ-27 operationalization in Table 2; no label or benchmark number is defined in terms of the paper's own outputs or conclusions. The benchmark compares model predictions against the human-annotated labels, which is standard supervised evaluation, and the qualitative error analysis in Section 4.3 is a post hoc reading of GPT-4o's 137 misclassifications rather than a claim derived from the annotation scheme. The author self-citations (Lee et al., 2023; Meng et al., 2024) are used to justify combining AQ-27 subscales and interview structure; those are design decisions, not the paper's predicted result, and they do not force any finding. The central claims are factual corpus properties (4,141 snippets from 684 participants, Cohen's kappa of 0.71) and model performance numbers, all verifiable from the released data and standard evaluation protocols. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no theoretical result is asserted by definition. Accordingly, no circular step is present.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central claim is a resource claim, so the load-bearing premises are the construct validity of the attribution-based interview and the reliability of the two-annotator labels. One fitted design parameter (follow-up length thresholds) shapes snippet composition. No new theoretical entities are introduced; the fictional character Avery is a standard vignette instrument, not an invented theoretical construct.

free parameters (1)
  • Follow-up question length thresholds = 25 and 150 characters
    The questioning protocol q(r) in Section 3.2.1 uses thresholds of 25 and 150 characters to decide whether participants receive one or two follow-up questions. These were determined through an 8-participant pilot study and specialist consultation, and they shape snippet length and composition across the corpus.
assumptions (5)
  • domain assumption Attribution theory (Kelley, 1967) and Corrigan et al.'s (2003) attribution model provide a valid decomposition of public mental-health stigma.
    Invoked in Section 3.1 as the theoretical basis for both interview questions and annotation labels. The central claim that the corpus captures fine-grained stigma depends on this theory being appropriate.
  • domain assumption The AQ-27 (Corrigan, 2012), originally a survey instrument, can be validly adapted to an open-ended chatbot interview without changing the measured construct.
    Section 3.1 states that the AQ-27 was adapted into seven core interview questions. If the adaptation changes what participants express, the labels do not measure the intended attributions.
  • domain assumption Chatbot-facilitated interviews with a fictional vignette elicit genuine stigmatizing attitudes rather than only socially desirable responses.
    Section 3.2.1 relies on scenario-embedded questions and randomized order to mitigate social-desirability bias. The validity of the corpus presupposes that participants answered candidly.
  • domain assumption Expert-guided annotation by two research assistants, with consensus discussion and a Cohen's kappa of 0.71, produces reliable ground-truth labels for the benchmark.
    Section 3.3 describes the annotation setup. The benchmark scores are interpreted against these labels. Two annotators and consensus resolution leave unmeasured subjectivity.
  • ad hoc to paper Excluding participants with immediate mental-health concerns (K6 screening) does not materially distort the stigma distribution the corpus intends to represent.
    Section 3.2.2 applies this inclusion criterion to safeguard vulnerable participants, but it also removes a segment of the population whose stigma responses might differ systematically.

how reviews work

0 comments
Cite this review

Pith. "Pith review of What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma." pith.science (2026). https://pith.science/paper/M7LDYUW7

@misc{pith2026250512727,
  author       = {Pith},
  title        = {Pith review of: What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/M7LDYUW7}},
  note         = {Machine review of arXiv:2505.12727}
}
read the original abstract

Mental-health stigma remains a pervasive social problem that hampers treatment-seeking and recovery. Existing resources for training neural models to finely classify such stigma are limited, relying primarily on social-media or synthetic data without theoretical underpinnings. To remedy this gap, we present an expert-annotated, theory-informed corpus of human-chatbot interviews, comprising 4,141 snippets from 684 participants with documented socio-cultural backgrounds. Our experiments benchmark state-of-the-art neural models and empirically unpack the challenges of stigma detection. This dataset can facilitate research on computationally detecting, neutralizing, and counteracting mental-health stigma. Our corpus is openly available at https://github.com/HanMeng2004/Mental-Health-Stigma-Interview-Corpus.

Figures

Figures reproduced from arXiv: 2505.12727 by the authors.

Figure 1
Figure 1. A sample interview snippet from our corpus [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Token-count distribution across attributions (left), the association between mental-health stigma and [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Chatbot interface for (a) vignette delivery and (b) an example interview question, showing a multi-turn [PITH_FULL_IMAGE:figures/full_fig_p018_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Screenshot of the annotation platform interface. [PITH_FULL_IMAGE:figures/full_fig_p021_4.png]
Figure 5
Figure 5. Figure 5: Annotation instructions for human annotators. [PITH_FULL_IMAGE:figures/full_fig_p023_5.png]
Figure 6
Figure 6. Figure 6: Heatmap showing the agreement between two human annotators. [PITH_FULL_IMAGE:figures/full_fig_p024_6.png]
Figure 8
Figure 8. Figure 8: Word-embedding comparison of 1,000 random samples each from [PITH_FULL_IMAGE:figures/full_fig_p025_8.png]
Figure 9
Figure 9. Figure 9: Visualization of BERT embeddings for correctly ( [PITH_FULL_IMAGE:figures/full_fig_p025_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

23 extracted references · 20 canonical work pages

  1. [1]

    36 Keywords: responsible, responsibility, weakness, stem from, blame, etc

    Responsibility: A common stigmatized thought is that people have control over and are responsible for their mental illness and related symptoms. 36 Keywords: responsible, responsibility, weakness, stem from, blame, etc

  2. [2]

    Keywords: worried, apprehensive, upset, unhappy, worrying, get along, distance, etc

    Social Distance: The constant stigmatized behavior is to stay away from people with mental illness. Keywords: worried, apprehensive, upset, unhappy, worrying, get along, distance, etc

  3. [3]

    Keywords: annoyed, angry, anger, irritated, lost temper, etc

    Anger: The stigmatized thought is to be irritated or annoyed because people are blamed for their mental illness. Keywords: annoyed, angry, anger, irritated, lost temper, etc

  4. [4]

    Keywords: lack, hinder, refuse, less inclined to, help, assist, etc

    Helping: The stigmatized behavior is withholding support towards them because of biased thoughts. Keywords: lack, hinder, refuse, less inclined to, help, assist, etc

  5. [5]

    Keywords: concern, pity, sympathy, tough, sad, lack, etc

    Pity: A common stigmatized thought is to be unsympathetic towards people with mental illness, and place blame on them for their mental health challenges. Keywords: concern, pity, sympathy, tough, sad, lack, etc

  6. [6]

    Keywords: separate, lonely, warrant, hospitalization, distance, treatment, respect, neighborhood, neighbor, socialization, companionship, etc

    Coercive Segregation: The stigmatized behavior is to send people with mental illness to institutions away from their community and force people with mental illness to participate in medication management or other treatments. Keywords: separate, lonely, warrant, hospitalization, distance, treatment, respect, neighborhood, neighbor, socialization, companion...

  7. [7]

    Keywords: violence, threatened, safe, danger, afraid, cautious, hurt, safety, frightened, scare, intimidate, threatening, anger, fear, yelling, etc

    Fear: The stigmatized thought is to believe that people with mental illness are not safe or feel frightened since people with mental illness are dangerous or unpredictable. Keywords: violence, threatened, safe, danger, afraid, cautious, hurt, safety, frightened, scare, intimidate, threatening, anger, fear, yelling, etc

  8. [8]

    Please carefully read the additional rules and adhere strictly to them when annotating the data:

    Non-stigmatizing: The text explicitly conveys non-stigmatizing content, displaying positive attitudes such as sympathy, assistance, friendliness, etc. Please carefully read the additional rules and adhere strictly to them when annotating the data:

Show all 23 references
  1. [9]

    If it doesn't have any stigmatized thoughts, choose H. ''' 38

  2. [15]

    If a message mentions both stigmatized and non-stigmatized statements, annotate it as Stigmatizing

  3. [16]

    People have control over and are responsible for their mental illness and related symptoms

    When discussing ‘responsibility,’ common stigmatized thoughts are: i. People have control over and are responsible for their mental illness and related symptoms. ii. They think their situation is a result of their actions. iii. They think it’s a personal weakness. iv. They can...

  4. [17]

    Preferring to stay away from people with mental illness

    When discussing ‘social distance,’ common stigmatized thoughts are: vi. Preferring to stay away from people with mental illness. vii. Associating mental illness with negative living situations or predicting future conflicts. 37 viii. Rejecting the idea of living with people wi...

  5. [18]

    Feeling irritated or annoyed with individuals because of their mental illness

    When discussing 'anger', common stigmatized thoughts are: x. Feeling irritated or annoyed with individuals because of their mental illness. xi. Feeling uncomfortable around people with mental illness. xii. Expressing ‘dislike’ for people with mental illness. xiii. Other statem...

  6. [19]

    Withholding support or assistance towards people with mental illness

    When discussing 'helping', common stigmatized thoughts are: xiv. Withholding support or assistance towards people with mental illness. xv. Offering indirect help, e.g., guidance, to people with mental illness. xvi. Offering conditional help to people with mental illness. xvii....

  7. [20]

    Feeling no concern or sympathy for individuals with mental illness

    When discussing 'pity', common stigmatized thoughts are: xix. Feeling no concern or sympathy for individuals with mental illness. xx. Other statements reflect the mean opinions. If a message contains the stigmatized thought mentioned in xix. - xx., choose E

  8. [21]

    Suggesting for sending people with mental illness to institutions away from their community

    When discussing 'coercive segregation', common stigmatized thoughts are: xxi. Suggesting for sending people with mental illness to institutions away from their community. xxii. Suggesting that people with mental illness need to take the medication or treatment, with or without...

  9. [22]

    Feeling frightened of people with mental illness

    When discussing 'fear', common stigmatized thoughts are: xxvi. Feeling frightened of people with mental illness. xxvii. Viewing people with mental illness as dangerous and unpredictable. xxviii. feeling extra cautious when interacting with people with mental illness. xxix. Ass...

  10. [1964]

    how" or “what

    – that language might be able to capture. B More Details about Data Collection: Chatbot-based Interview B.1 Vignettes The clinical version appears below: Avery is employed by a company, and in their spare time, they are dedicated to lifelong learning, doing extensive reading a...

  11. [2020]

    i hear you, i feel you

    Do models of mental health based on social media data generalize? In Findings of the Associa- tion for Computational Linguistics: EMNLP 2020 , pages 3774–3788, Online. Association for Computa- tional Linguistics. Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, D...

  12. [2021]

    HateCheck: Functional tests for hate speech detection models. In Proceedings of the 59th An- nual Meeting of the Association for Computational Linguistics and the 11th International Joint Confer- ence on Natural Language Processing (Volume 1: Long Papers), pages 41–58, Online....

  13. [2022]

    In Proceedings of the 2022 Con- ference on Empirical Methods in Natural Language Processing, pages 2152–2170, Abu Dhabi, United Arab Emirates

    Gendered mental health stigma in masked language models. In Proceedings of the 2022 Con- ference on Empirical Methods in Natural Language Processing, pages 2152–2170, Abu Dhabi, United Arab Emirates. Association for Computational Lin- guistics. Bruce G. Link, Francis T. Cullen...

  14. [2023]

    In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 14523–14530

    Everyone’s voice matters: Quantifying anno- tation disagreement using demographic information. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 14523–14530. Zeerak Waseem and Dirk Hovy. 2016. Hateful symbols or hateful people? predictive featu...

  15. [2024]

    social prim- ing

    Exploring the relationship between intrin- sic stigma in masked language models and train- ing data using the stereotype content model. In Proceedings of the Fifth Workshop on Resources and ProcessIng of linguistic, para-linguistic and extra-linguistic Data from people with va...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.