{"id":"e692492b-10d2-4a6c-b552-de9e7bbb76c4","arxiv_id":"2505.12727","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper introduces the first large-scale interview-based, theory-grounded expert-annotated corpus for fine-grained mental-health stigma detection, plus model benchmarks showing LLMs still struggle with subtle stigma.","lead":"The authors built and released an expert-annotated corpus of 4,141 chatbot-interview snippets from 684 people, labeling mental-health stigma into seven theory-derived categories. The paper also benchmarks large language models on this data and reports that even GPT-4o misses subtle, often well-intentioned stigmatizing language.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The adapted AQ-27 interview lacks construct validation, and the annotation rule that any mixed response is stigmatizing likely over-attributes stigma, undermining the corpus labels.","rationale":"The paper is a careful and transparent dataset contribution, and the reader's conditional verdict is appropriate. My stress-test identifies construct validity as the most load-bearing link in the argument. The authors adapted a standardized self-report questionnaire into an open-ended chatbot interview without validating that the open-ended responses correspond to the same theoretical attributions. The annotation rule that any mixed statement is stigmatizing is a strong, contestable decision that likely inflates stigma prevalence; Figure 1 shows a participant explicitly agreeing to rent to Avery yet being labeled stigmatizing. Because the benchmark results and the qualitative conclusions are all measured against these labels, label validity is central rather than peripheral. The proposed validation study—correlating chatbot-based labels with AQ-27 scores on the same participants—would directly test the construct assumption, while a second annotator check would test label robustness. The paper's own limitations section acknowledges perspectivism and the need for annotator-level labels, which is honest, but the absence of any empirical validation means the central claim is currently conditional. I therefore do not change the reader's verdict; the paper should be accepted only with the requirement that the authors either provide such validation or soften the claim of theory-grounded measurement.","tokens_in":33754,"tokens_out":8678,"duration_ms":94008,"concrete_test":"Run a validation study with a new sample of ~100 participants: each completes the chatbot interview and the original AQ-27 (or a validated short form) in counterbalanced order. For each of the seven attributions, compute Spearman correlations between the AQ-27 subscale score and the binary stigma label derived from the annotation protocol. If the average correlation is below 0.3, the adapted interview does not capture the same construct. In parallel, have two additional annotators from a different cultural background independently re-annotate a random sample of 200 snippets using the released codebook; if Cohen's kappa with the original labels is below 0.5, the labels are not robust to annotator background and the mixed-response rule should be re-examined.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that this corpus measures Corrigan et al.'s mental-health stigma construct rests on the operationalization in Section 3.1, where AQ-27 Likert items are converted into seven open-ended chatbot questions, and on the annotation protocol in Section 3.3. No concurrent or discriminant validation against the AQ-27 or any established stigma scale is reported; the only support is face validity (Table 2). This matters because the annotation codebook contains the rule 'If a message mentions both stigmatized and non-stigmatized statements, annotate it as Stigmatizing' (Appendix H.3, rule 1). Applied to hedged conversational responses, this rule labels participants who ultimately express acceptance as stigmatizing based on an initial hesitation. Figure 1 is a clear case: the participant says 'if I were a landlord looking for a tenant I would rent to them' but is labeled Stigmatizing (Social Distance) because of the preceding 'I might have to think more about it.' Whether this response is stigmatizing is contestable. Since these labels serve as ground truth for the benchmarks (Table 4) and for the qualitative claim in Section 4.3 that neural models are insufficient, systematic over-attribution would inflate apparent model failures and distort the corpus's theoretical grounding. The reliance on two annotators, both Asian and in their twenties, with a moderate kappa of 0.71, and the absence of a cross-cultural annotation check, further weaken the epistemic status of the labels. If the adapted measure is not the same construct, the paper's central claim of a theory-grounded stigma corpus is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces MHS TIGMA INTERVIEW, an expert-annotated, theory-informed corpus for mental-health stigma detection, comprising 4,141 interview snippets from 684 participants who interacted with a chatbot about a vignette character with depression. Seven interview questions adapt items from the Attribution Questionnaire-27 (AQ-27) to probe responsibility, social distance, anger, helping, pity, coercive segregation, and fear. Two trained annotators labeled each snippet with one of seven stigma attributions or a non-stigmatizing category, with checkpoint kappas and a final Cohen's kappa of 0.71. The authors benchmark fine-tuned RoBERTa and several large language models under zero-shot, one-shot, and full-codebook prompting, and qualitatively analyze 137 GPT-4o misclassifications to characterize subtle stigmatizing language. The paper claims to provide the first large-scale, open-source mental-health stigma interview dataset with theoretical grounding and documented socio-cultural backgrounds, and concludes that neural models alone remain insufficient for stigma detection.","tokens_in":33938,"tokens_out":5879,"duration_ms":63015,"significance":"If the corpus labels are valid, this is a genuinely useful resource. It addresses a real gap left by social-media and synthetic corpora, uses a recognized psychological framework, and ships with unusually thorough documentation: IRB approval, consent procedures, iterative codebook development, checkpoint agreement, an agreement matrix, annotator feedback, and access-controlled release. The benchmark experiments and the qualitative error analysis are valuable starting points for future work. The main risk is construct validity: the adapted interview protocol and the annotation rules may over-attribute stigma, and this would propagate into every benchmark number and qualitative conclusion. The paper deserves serious consideration, but the labeling assumptions need to be tested or explicitly softened before the strongest claims can be accepted.","major_comments":[{"comment":"The operationalization of the AQ-27 as seven open-ended chatbot questions is supported only by face validity; no concurrent or discriminant validation against the original AQ-27 or another stigma scale is reported. This is load-bearing because the corpus's claim to be 'theory-grounded' and the benchmark conclusions in Section 4.2 both rest on the assumption that the chatbot questions elicit the same construct as the AQ-27. I would ask for a small validation subsample (e.g., administering the AQ-27 to a subset of participants) or, at minimum, an explicit reframing of the claim to 'theory-informed' and a discussion of the validation gap. The Limitations section does not currently acknowledge this missing validation.","section":"3.1, Table 2, Limitations"},{"comment":"Annotation rule 1 in the codebook ('If a message mentions both stigmatized and non-stigmatized statements, annotate it as Stigmatizing') systematically resolves mixed conversational responses toward stigma. Figure 1 illustrates the risk: the participant ultimately says they would rent to Avery, yet the snippet is labeled Stigmatizing (Social Distance) because of the preceding hesitation. Because these labels are the ground truth for Table 4 and the qualitative claim in Section 4.3, this rule can overstate stigma prevalence and inflate apparent model failures. I recommend either refining the rule for hedged responses or reporting sensitivity analyses with a more conservative rule.","section":"Appendix H.3, Figure 1"},{"comment":"The final labels rest on two annotators with similar demographic backgrounds (both Asian and in their twenties), with a moderate Cohen's kappa of 0.71, and disagreements are resolved by consensus. The Limitations section candidly acknowledges annotator subjectivity and promises annotator-level labels only in a future v2.0, but the released corpus and all benchmark results use the consensus labels. For a construct as culturally and socially variable as stigma, I would like to see at least a disagreement analysis or an external audit by additional annotators from different backgrounds before treating these labels as reliable ground truth.","section":"3.3, Limitations"},{"comment":"The interview format and participant screening introduce assumptions that are not validated. Excluding participants with immediate mental-health concerns via the K6, and asking about willingness to rent, help, or hospitalize a fictional character through a chatbot, may suppress or redirect the very attitudes the corpus aims to measure. Section 3.2.1 cites scenario-embedded questions and randomized order to mitigate social-desirability bias, but no evidence is provided that the chatbot setting elicits candid stigmatizing attitudes. This is connected to the construct-validity concern above and should be addressed directly, for example by comparing chatbot interview responses with a standard survey administration in a pilot study.","section":"3.2.1, 3.2.2"}],"minor_comments":[{"comment":"Demographic documentation is based on 555 of 684 participants (81.1%); the abstract's phrase 'documented socio-cultural backgrounds' should be qualified to avoid overstating coverage.","section":"Abstract, Section 3.4"},{"comment":"The snippet in Figure 1 appears with three labels (Stigmatizing (Responsibility), Stigmatizing (Social Distance), Non-stigmatizing), while Section 3.3 describes a single-label task; please clarify whether this is a multi-label illustration or an error in the figure.","section":"Figure 1"},{"comment":"The entry for Roesler et al. uses '1' in the Theory-Grounded column with no footnote; use checkmarks or explain the partial mark.","section":"Table 1"},{"comment":"The 25- and 150-character thresholds in the follow-up question policy are justified only by an 8-participant pilot; a sentence on the robustness of these thresholds would be helpful.","section":"Equation (1), Section 3.2.1"},{"comment":"The keyword 'distance' appears in both Social Distance and Coercive Segregation definitions, which may confuse both annotators and prompted models; consider disambiguating the two sets of keywords.","section":"Appendix H.3"},{"comment":"GPT-4o was evaluated with a single run and a single temperature (0.2) due to budget constraints; the reported F1 differences should be interpreted with this uncertainty in mind, and a caveat in the text would help.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The dataset paper is a good fit for the journal and the resource is potentially valuable. The main issue is the missing construct validation, which I believe is addressable with a focused validation study or a careful reframing of the claims; I do not see a need to reject the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful corpus paper, and the corpus is the contribution. The authors have done something no one else has: collected 4,141 interview snippets from 684 participants via a chatbot, grounded the annotation in Corrigan's attribution model, and released socio-cultural metadata. The annotation process is transparent—IRB, expert review, pilot thresholds, iterative codebook, checkpoint kappas, agreement matrix. That alone merits taking it seriously.\n\nThe real soft spot is construct validity. The authors adapt AQ-27 Likert items into open-ended chatbot questions but never validate the adaptation against the original scale or any other stigma measure. The mapping in Table 2 is face-valid, but face validity doesn't establish that an open-ended conversation about a fictional character taps the same attributional judgments as a standardized survey. The stress-test note is right to flag this.\n\nMore concretely, the codebook rule that any message containing both stigmatizing and non-stigmatizing statements must be labeled Stigmatizing is likely to over-attribute. Figure 1 is a clear case: the participant says they would rent to Avery, but the \"I might have to think more about it\" gets labeled Social Distance. That is contestable. If a nontrivial share of the corpus is labeled this way, the benchmark numbers and the qualitative \"models fail\" claim inherit the bias. This is not fatal—the rule may be defensible as a conservative stance for a detection task—but it needs quantification, justification, and ideally a sensitivity analysis with a \"mixed\" label.\n\nThe other issues are minor. Two annotators is thin for a subjective construct, though the authors acknowledge this and promise annotator-level labels in v2.0. Benchmark variance is under-reported (GPT-4o is a single run). The demographic associations in Section 3.4 are called preliminary, so the missing significance tests are less concerning. Also, someone should verify the GitHub repository actually matches the paper; that is the artifact.\n\nWho is this for? People building stigma-detection models, HCI researchers studying chatbot interviews, and computational social scientists wanting demographically documented text. It deserves a serious referee. I'd send it to review with a request for construct validation, a sensitivity analysis on the mixed-response rule, and error bars.","headline":"A genuinely new and useful interview-based stigma corpus, but the theory-grounded label claim needs construct validation and the mixed-response annotation rule likely over-attributes stigma.","tokens_in":34584,"tokens_out":2849,"would_cite":true,"duration_ms":32056,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a new expert-annotated interview corpus, built on attribution theory, lets researchers detect fine-grained mental-health stigma in conversation and shows that state-of-the-art language models still miss its most subtle…","keywords":["mental-health stigma","interview corpus","attribution theory","chatbot interview","multi-label annotation","stigma detection benchmark","expert annotation","large language models"],"falsifier":"Administer the original AQ-27 to the same 684 participants shortly after their chatbot interview and compute person-level correlation between survey scores and the stigma labels assigned to their interview snippets; if the correlation is negligible, the interview adaptation is not measuring the same construct. A complementary check would compare chatbot-elicited responses with anonymous written self-reports on the same vignette to detect social-desirability suppression.","tokens_in":33453,"feed_emoji":"💬","tokens_out":7609,"duration_ms":73057,"temperature":0.7,"pith_summary":"The paper sets out to give computational research on mental-health stigma a data source it has lacked: a large, open, theory-grounded corpus of how ordinary people actually talk about stigma in conversation. It introduces 4,141 expert-annotated transcript snippets drawn from 684 participants in structured human-chatbot interviews, each snippet labeled as non-stigmatizing or as one of seven attribution types (responsibility, social distance, anger, helping, pity, coercive segregation, fear) derived from a standard attribution model of public stigma. The authors argue that the corpus captures veiled, socially acceptable forms of stigma that social-media and synthetic corpora miss, and they benchmark open models and GPT-4o to show that state-of-the-art neural classifiers still misread these subtleties. If the corpus does what it claims, it supplies a shared benchmark for detecting, neutralizing, and counteracting mental-health stigma in language.","feed_headline":"GPT-4o and top LLMs miss subtle mental-health stigma","feed_subtitle":"A 4,141-snippet expert-labeled interview corpus reveals veiled stigma that state-of-the-art models overlook.","key_machinery":"The load-bearing object is the MHS TIGMA INTERVIEW corpus itself, a collection of 4,141 expert-annotated snippets from chatbot interviews, and the attribution model that shapes it. The annotation scheme condenses the attribution model and AQ-27 into seven stigma attributions, each linked to one interview question embedded in a vignette about the fictional character Avery. The interview protocol's role is to elicit candid reactions with scenario-embedded questions, randomized order, follow-up prompts, and neutral self-disclosure from the chatbot, and the benchmark experiments then use the corpus to test how well models can recover the human labels.","core_discovery":"The central claim is that mental-health stigma can be operationalized into seven measurable attribution types and that a carefully elicited interview corpus can surface real stigmatizing talk that existing datasets and models miss. The authors adapt the Attribution Questionnaire-27 into seven scenario-based chatbot questions about a fictional person with depression, collect responses from 684 screened English-speaking adults, and have two expert annotators label each snippet under a codebook informed by the attribution model, reaching 0.71 Cohen's kappa. The resulting corpus labels 46% of snippets as stigmatizing, with responsibility and social distance the most common types, and with toxicity lower than in hate-speech benchmarks, evidence that stigma here appears as veiled, normalized, or well-intentioned language. Benchmarks on eight-way classification show that larger instruction-tuned models improve with a full codebook but still fall short, with GPT-4o's best macro F1 around 0.757, and a qualitative analysis of GPT-4o's 137 misclassified snippets identifies distancing language, misuse of psychiatric terms, coercive phrasing, differential support, patronization, and minimization as the recurring failure modes.","pith_inferences":["Testable extension: give the same participants the original AQ-27 survey and correlate their scores with the stigma labels from their interview snippets; a weak correlation would suggest the chatbot adaptation changes the construct being measured.","Because each interview snippet is tied to the attribution-specific question that elicited it, prompt-ablating that question text would test whether model errors come from the wording of the question rather than from the participant's stigma.","The same interview and annotation protocol could be run in non-Western languages and cultures; the current corpus, sampled mostly from Western English-speaking participants, cannot show whether the seven attribution types structure stigma expression everywhere."],"forward_implications":["Systems can be trained to classify stigma by its underlying attribution (blame, distance, anger, withholding help, pity, forced treatment, fear) instead of treating it as a single toxic category.","Researchers get a benchmark where the full codebook given to human annotators measurably improves LLM performance, so gains from better instructions can be tracked against expert labels.","The documented demographic and geographic metadata make it possible to study how stigma expression varies with gender, country, education, and personal exposure to mental illness.","The corpus can be used to have models role-play interviewees and compare their generated responses with the human ones, exposing whether models internalize and perpetuate the same stigmatizing attributions."],"supporting_citations":[{"why":"Supplies the attribution model of public stigma that defines the corpus's cognitive, emotional, and behavioral components and its seven attribution types.","marker":"Corrigan et al., 2003"},{"why":"Provides the AQ-27 survey items that were adapted into the seven scenario-based interview questions.","marker":"Corrigan, 2012"},{"why":"Provides the K6 screening instrument used to exclude participants with pressing mental-health concerns.","marker":"Kessler et al., 2003"},{"why":"Informs the chatbot-based social-contact interview design, including the vignette about a person with depression.","marker":"Lee et al., 2023"},{"why":"Is the existing synthetic mental-health stigma corpus that serves as the main comparison point in the related-work discussion.","marker":"Choey, 2023"},{"why":"Supplies microaggression data from Tumblr used as a comparison in the semantic and toxicity analyses.","marker":"Breitfeller et al., 2019"},{"why":"Supplies implicit hate speech data from Twitter used as a comparison benchmark in the toxicity and embedding analyses.","marker":"ElSherief et al., 2021"},{"why":"Provides the POTATO annotation platform on which the two annotators labeled the interview snippets.","marker":"Pei et al., 2022"},{"why":"Supplies the RoBERTa-base model that is fine-tuned as the encoder-only benchmark for stigma detection.","marker":"Liu et al., 2019b"}],"fun_headline_variants":["LLMs stumble on veiled mental-health stigma","New corpus exposes stigma that AI models miss","Mental-health stigma: even GPT-4o can't spot it all","Subtle mental-health stigma: benchmark reveals AI gap","Chatbot interviews reveal hidden stigma AI overlooks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that adapting the Attribution Questionnaire-27 into seven chatbot interview questions about a fictional person measures the same mental-health stigma construct as the original survey, and that participants answered candidly rather than editing their responses to appear socially acceptable.","fun_headline_variants_meta":{"raw":{"variants":["LLMs stumble on veiled mental-health stigma","New corpus exposes stigma that AI models miss","Mental-health stigma: even GPT-4o can't spot it all","Subtle mental-health stigma: benchmark reveals AI gap","Chatbot interviews reveal hidden stigma AI overlooks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1290,"prompt_tokens":905,"completion_tokens":385,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":309}},"tokens_in":521,"tokens_out":385,"duration_ms":4090,"temperature":1.0,"reasoning_tokens":309,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:27:11.823785+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Administer the original AQ-27 to the same 684 participants shortly after their chatbot interview and compute person-level correlation between survey scores and the stigma labels assigned to their interview snippets; if the correlation is negligible, the interview adaptation is not measuring the same construct. A complementary check would compare chatbot-elicited responses with anonymous written self-reports on the same vignette to detect social-desirability suppression.","supporting_citations":[],"review_version":1}