{"id":"a1886b25-43da-437d-ba9c-3a744d722049","arxiv_id":"2506.04043","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Persona- and emotion-guided prompts make LLM counter-narratives more empathetic and readable, but generated responses remain verbose, college-level, and sometimes classified as hateful.","lead":"Researchers tested three AI chatbots that write responses to online hate speech, using different 'persona' prompts to see which replies were kinder, easier to read, and safer. The results reveal trade-offs between friendliness, readability, and safety that matter for anyone building AI moderation tools.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central 'empathetic' finding is supported only by emotion classifiers that the paper itself shows are driven by surface lexical cues, and the NGO-Emotion prompt's explicit 'compassionate' instruction creates a likely prompt-classifier confound.","rationale":"The load-bearing concern is the validity of the emotion classifiers used to support the 'empathetic' conclusion. The reader's weakest_assumption identified the same issue, and my analysis agrees. The paper's own Section 4.4 provides internal evidence of classifier unreliability, and the NGO-Emotion prompt's explicit 'compassionate' wording creates a direct lexical confound: the generation prompt and the evaluation categories share vocabulary. This is a correctness risk, not a disagreement with consensus; the paper could be right, but the reported evidence does not establish the claim. The readability and verbosity findings are measured with standard metrics and are less affected, so they do not carry the same risk. I considered whether the lack of significance testing is more load-bearing, but a significant result on an invalid measure would still not establish empathy; the measurement validity issue is upstream. A single lexical-masking experiment can settle whether the 'empathetic' advantage is an artifact. Because the paper is a descriptive evaluation with other useful contributions and the authors could add this robustness check, I do not recommend changing the conditional verdict; I recommend keeping CONDITIONAL (UNCHANGED).","tokens_in":18360,"tokens_out":5110,"duration_ms":50892,"concrete_test":"Take a random sample of 200 NGO-Emotion and 200 Vanilla responses per model (approximately 1,200 total). Mask all tokens matching a compassion-related lemma list (compassion*, care*, kind*, empat*, understand*, support*, respect*, lov*, etc.) in both conditions, then re-run the RoBERTa and Mistral emotion classifiers on the masked texts. If the NGO-Emotion advantage in 'caring'/'approval' disappears or reverses, the empathy finding is a lexical artifact of the prompt. If the advantage persists after masking, the concern is substantially mitigated and a small human annotation study would be the next check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central abstract claim is that 'emotionally guided prompts yield more empathetic and readable responses.' The 'empathetic' component rests entirely on automated emotion labels from RoBERTa-on-GoEmotions and Mistral-7B (Sections 3.3 and 4.4). The paper itself documents that these classifiers are not reliable for this purpose: Section 4.4 reports RoBERTa 'appeared to over-rely on certain lexical cues, such as fun, happy, party, celebrate, and enjoy,' and Cohere responses labeled 'love' were 'largely driven by surface-level lexical indicators, particularly the frequent inclusion of the word love.' Moreover, the NGO-Emotion prompt (Table 8) explicitly instructs the model to 'generate a concise, well reasoned, and compassionate counter-narrative.' Consequently, generated responses are likely to contain compassion-related vocabulary (e.g., 'compassion,' 'care,' 'understand'), which the GoEmotions classifier maps to 'caring' and 'approval,' and Mistral tends to assign the same labels. The observed increase in 'caring' under NGO-Emotion may therefore reflect lexical adherence to the instruction rather than genuine empathetic tone. The paper provides no human evaluation, no inter-annotator agreement, no significance tests, and no robustness analysis controlling for these surface cues. Since the 'empathetic' advantage is the paper's headline contribution, this confound is load-bearing: if it holds, the abstract's main claim is not supported by the reported evidence, even though the readability and verbosity findings are less affected.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multi-faceted evaluation framework for LLM-generated counter-narratives (CNs), comparing GPT-4o-Mini, Cohere CommandR-7B, and LLaMA 3.1-70B under three prompting strategies (Vanilla, NGO-Persona, NGO-Emotion) on the MT-Conan and HatEval datasets. It measures verbosity, readability (Flesch metrics), refusal rates, sentiment/emotion via DistilBERT, RoBERTa, and Mistral, and hatefulness via MetaHateBERT. The main reported findings are that LLM-generated CNs are verbose and require college-level literacy; NGO-Emotion prompts produce more verbose, more readable, and (claimed) more empathetic responses; and Cohere is cost-effective but behaviorally inconsistent. The authors also highlight limitations of the automated classifiers and hatefulness scores.","tokens_in":18696,"tokens_out":5250,"duration_ms":47375,"significance":"The proposed framework addresses a real gap by jointly evaluating persona framing, affective tone, readability, and ethical risk in CN generation, and the cross-model/cross-dataset design is a useful contribution. The verbosity and readability findings rest on deterministic, standard metrics and are credible; the comparison against human-authored CNs is a valuable baseline. The release of code, data, and responses supports reproducibility. However, the central claim that emotionally guided prompts yield 'more empathetic' responses is not supported by the evidence, because the emotion classifiers are shown in the paper itself to be confounded by surface lexical cues, and the NGO-Emotion prompt explicitly instructs 'compassionate' language. The safety-related hatefulness findings are similarly weakened by the acknowledged false-positive problem of MetaHateBERT. If the empathy and safety claims are tempered or re-supported with human evaluation, the framework and readability findings would still be of interest to the CN-generation community.","major_comments":[{"comment":"The 'empathetic' advantage claimed for NGO-Emotion prompts is confounded. Table 8's NGO-Emotion prompt explicitly instructs the model to generate a 'compassionate' counter-narrative 'with empathy'; Appendix A.2 states that Mistral maps 'empathy, compassion, and understanding' to the 'caring' label; and Section 4.4 documents that both RoBERTa and Mistral emotion labels are driven by surface lexical cues (e.g., 'fun', 'happy', 'party', 'love'). The observed increase in 'caring'/positive emotions under NGO-Emotion is therefore expected from lexical adherence to the prompt rather than evidence of more empathetic responses. No human evaluation, inter-annotator agreement, or robustness analysis controlling for prompt-specific vocabulary is provided. Since the abstract's headline claim is that 'emotionally guided prompts yield more empathetic and readable responses,' this confound is load-bearing. I recommend either removing the 'empathetic' claim from the central conclusions or supporting it with human ratings and/or a control analysis that strips or neutralizes prompt-specific terms.","section":"§4.4, Table 8, Appendix A.2"},{"comment":"The hatefulness results are presented as 'Hatefulness Scores' in Table 4, yet the paper itself provides strong evidence that MetaHateBERT misclassifies CNs that merely reference or condemn hateful content (Section 4.5, Table 12). For example, the CN 'Jews are a diverse community...' receives a hatefulness score of 0.80 despite being clearly non-hateful. Reporting these raw classifier outputs without correction or human validation overstates the safety risk and weakens the abstract's 'concerns surrounding safety.' The authors should either present these as raw classifier scores with explicit false-positive caveats, or add a human-validated subset to estimate the actual hatefulness rate.","section":"§4.5, Table 4, Section 5 (Limitation of hate speech classification systems)"},{"comment":"All comparative claims are based on descriptive statistics without significance tests or confidence intervals. For instance, Section 4.1 states that NGO-Emotion prompts 'yield the most readable outputs' and Section 4.3 claims that 'NGO-Emotion significantly enhances positive sentiment,' but no inferential tests are reported. Given the large sample sizes (n=2000 and n=5003), even small differences may be statistically significant, or conversely, observed differences could be within sampling noise. I strongly encourage adding per-comparison confidence intervals and tests (e.g., bootstrap or paired tests) or tempering the language to describe only descriptive patterns.","section":"Results (Tables 1-4), §4.1, §4.3"}],"minor_comments":[{"comment":"The text says MT-Conan's original text has a mean word count of 13.6, but Table 1 reports 13.2; please make the numbers consistent.","section":"§4.1 and Table 1"},{"comment":"The statement that GPT-4o-mini has '20 billion' parameters is not supported by the cited GPT-4o system card, which does not disclose the parameter count; please rephrase to 'a smaller model' or remove the specific number.","section":"§5 (Cost vs Capability)"},{"comment":"The layout of Table 1's header ('Data Source Persona Dataset') is confusing; reorganize so the dataset columns are clearly aligned with the model/persona rows.","section":"Table 1 header"},{"comment":"There are grammatical errors: 'The command models was initially used' should be 'The command model was initially used,' and 'we decided to use this models' should be 'we decided to use these models.'","section":"Appendix B.1"},{"comment":"The title 'The GoEmotion Dataset' should be 'The GoEmotions Dataset' for consistency with the rest of the paper and the dataset's official name.","section":"Appendix A.2"},{"comment":"The dataset name is inconsistently spelled as both 'HateEval' (e.g., Section 4.1) and 'HatEval' (most other places); please use one spelling consistently.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid empirical evaluation with good reproducibility practices, but the central 'empathetic' finding is overinterpreted given the classifier confounds documented within the paper itself. The revision path is clear: either add human evaluation or reframe the claims to describe emotion-classifier outputs rather than empathy. The hatefulness results also need to be reported with appropriate caveats. These issues are within the scope of a major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on counterspeech or LLM evaluation. The paper's real contribution is a multi-dimensional evaluation framework — persona framing, verbosity/readability, affective tone, and hatefulness — applied across three models and two datasets, with a public artifact. The readability and verbosity results are the strongest part: deterministic metrics, a human-authored baseline (Flesch Reading Ease 59.6, grade 8.7), and a clear finding that LLM counterspeech mostly requires college-level literacy. The hatefulness analysis is also handled honestly; the authors flag likely false positives from MetaHateBERT and show examples where counterspeech is misclassified because it quotes the hateful text. That is the kind of self-aware evaluation work we need more of.\n\nThe soft spot is the headline. The claim that NGO-Emotion prompts yield \"more empathetic\" responses depends entirely on automated emotion classifiers — RoBERTa-on-GoEmotions and Mistral-7B. The paper itself demonstrates these classifiers over-rely on surface lexical cues (\"fun,\" \"happy,\" \"love\"), and Appendix A.2 explicitly maps \"empathy,\" \"compassion,\" and \"understanding\" to \"caring.\" The NGO-Emotion prompt tells the model to generate a \"compassionate counter-narrative.\" So the observed increase in \"caring\" under that condition is almost certainly the model following the instruction and the classifier matching the resulting vocabulary. That is a prompt–classifier confound, and it is load-bearing: without it, the abstract's main positive claim is unsupported. The readability finding survives because it uses deterministic Flesch scores; the verbosity finding survives for the same reason.\n\nAlso missing: no human evaluation, no inter-annotator agreement, no significance tests, no error bars, and the temperature is fixed but no seeds are reported. These are not fatal for the descriptive parts but they matter for any claim about empathy.\n\nWho is this for? Practitioners building counterspeech pipelines and researchers designing LLM evaluation studies. It deserves a serious referee, but the authors should be sent back to either add a human empathy evaluation or reframe the contribution around accessibility, verbosity, and classifier limitations. I would not cite the empathy claim as established.","headline":"Useful evaluation framework and credible readability/verbosity findings, but the headline empathy claim rests on emotion classifiers the paper itself shows are lexically biased, and the NGO-Emotion prompt's 'compassionate' instruction creates a direct confound.","tokens_in":19195,"tokens_out":1788,"would_cite":false,"duration_ms":20253,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Emotion-guided prompts make AI counter-speech more readable and empathetic, but not safer.","keywords":["counter-narratives","hate speech","large language models","emotion-guided prompting","readability","persona","affective computing","content moderation"],"falsifier":"A human annotation study in which native speakers rate counter-narratives for perceived empathy, warmth, and readability without knowing which prompt produced them would settle whether the NGO-Emotion advantage is real. If human ratings do not show a consistent advantage for the NGO-Emotion condition over NGO-Persona, the paper's central claim would be falsified. Additionally, a simple lexical probe--removing or varying words like 'love', 'happy', 'joy' from the generated texts--would reveal whether the classifier's 'empathetic' labels track those surface cues.","tokens_in":18193,"feed_emoji":"💬","tokens_out":1903,"duration_ms":19034,"temperature":0.7,"pith_summary":"This paper evaluates whether large language models can generate counter-narratives to hate speech that are not just factually sound but also accessible and emotionally appropriate. It compares three prompting strategies--vanilla, NGO-persona, and NGO-persona with explicit emotional guidance--across three LLMs and two hate-speech datasets. The central finding is that emotionally guided prompts produce responses that are more verbose, yet also more empathetic and more readable, suggesting that emotional framing can improve the quality of AI-generated counter-speech. However, the paper also finds that LLM-generated responses generally require college-level literacy, and that automated emotion classifiers may mistake surface-level words like 'love' for genuine empathy, so the safety and effectiveness of these responses remain in question.","feed_headline":"Emotion prompts make AI counter-speech more readable, study finds","feed_subtitle":"But the same prompts also make responses more verbose, and automated empathy scores may be faked by words like 'love'.","key_machinery":"The central mechanism is the multi-faceted evaluation framework that jointly measures persona framing, verbosity and readability, affective tone, and ethical robustness. For affect, the paper uses two classifiers: a RoBERTa model fine-tuned on GoEmotions and the Mistral-7B-Instruct LLM; for hatefulness, it uses MetaHateBERT. The key comparison is across three prompting strategies: Vanilla (no persona), NGO-Persona (explicit NGO worker role), and NGO-Emotion (the NGO persona plus explicit instruction to be compassionate). This framework lets the paper attribute changes in output quality to the emotional guidance in the prompt, while also exposing the classifiers' own biases.","core_discovery":"The paper's central claim is that prompting an LLM to adopt the persona of a compassionate NGO worker (the NGO-Emotion condition) consistently yields counter-narratives that are more verbose, more empathetic, and paradoxically more readable than either vanilla or NGO-persona prompting. This holds across GPT-4o-Mini, Cohere's CommandR-7B, and Meta's LLaMA 3.1-70B on both the MT-Conan and HatEval datasets. The paper also establishes that LLM-generated counter-narratives are generally harder to read than human-authored ones, with most requiring college-level literacy, and that smaller models like Cohere can produce more accessible text than larger ones, challenging the assumption that bigger models are always better. At the same time, the study finds that automated measures of emotion and hatefulness are unreliable: classifiers often rely on shallow lexical cues, and hate-speech detectors frequently misclassify counter-narratives that merely quote or condemn hateful language as hateful themselves.","pith_inferences":["A direct human study that rates the empathy and readability of the same counter-narratives, without exposing annotators to the prompting condition, would test whether the reported NGO-Emotion advantage is real or an artifact of the automated measures.","Because the classifiers' lexical biases are documented, one could design a simple probe: generate counter-narratives that explicitly avoid words like 'love', 'happy', and 'joy' and see whether the emotion labels shift, which would reveal how much of the 'empathetic' label is driven by word choice rather than meaning.","The finding that human-authored counter-narratives sit at an 8th-grade reading level suggests a concrete design target: LLMs could be prompted to match that level, and if they cannot, the implication is that human revision will remain essential for accessible moderation.","The paper's focus on English and on two target groups (immigrants and women) leaves open whether emotionally guided prompting would behave similarly for other languages or for hate speech targeting other groups, a natural extension for testing the generality of the claimed effect."],"forward_implications":["If emotionally guided prompts genuinely improve readability and perceived empathy, then prompt design alone can make AI counter-speech more accessible to people with lower literacy levels, which is a concrete step toward inclusive moderation.","The finding that smaller models (Cohere) can beat larger ones (LLaMA 70B) on readability suggests that model size is not the only driver of quality, and that deployment choices should weigh cost and capability jointly.","The inverse relationship between verbosity and readability implies that efforts to make outputs shorter may inadvertently make them harder to understand, so accessibility targets need explicit optimization.","The unreliability of automated emotion and hatefulness classifiers means that any evaluation of counter-speech quality that relies solely on such metrics may be misleading, and human evaluation remains necessary.","The paper's demonstration of classifier artifacts--e.g., 'love' being triggered by surface-level words--cautions against using automatic labels as ground truth in moderation pipelines."],"supporting_citations":[{"why":"Provides the MT-Conan dataset of hate speech and human-authored counter-narratives used as the primary benchmark.","marker":"Fanton et al. (2021)"},{"why":"Provides the HatEval dataset of real-world hate speech against women and immigrants, used as the secondary, more explicit dataset.","marker":"Basile et al. (2019)"},{"why":"Supplies the GoEmotions dataset and its 27 emotion categories that ground the RoBERTa emotion classifier's labels.","marker":"Demszky et al. (2020)"},{"why":"Contributes the MetaHateBERT model and the methodology for measuring hatefulness scores, which the paper applies to its generated counter-narratives.","marker":"Piot et al. (2024)"},{"why":"Establishes that LLMs can generate toxic outputs, motivating the paper's hatefulness evaluation and providing the baseline claim it investigates.","marker":"Piot and Parapar (2024)"},{"why":"Provides the vanilla prompting approach and an earlier analysis of GPT-generated counter-narratives on MT-Conan, which this paper extends with persona and emotion conditions.","marker":"Vallecillo Rodríguez et al. (2024)"}],"fun_headline_variants":["Emotion prompts boost empathy but also verbosity in AI counter-speech","LLM counter-narratives need college-level reading, study says","Smaller AI models can write more accessible hate-speech rebuttals","Hate detectors often misclassify counter-narratives that quote hate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that NGO-Emotion prompts produce more empathetic responses rests on the assumption that the automated emotion classifiers (RoBERTa and Mistral-7B) genuinely measure empathy rather than just detecting surface-level positive words, and the paper itself shows these classifiers often rely on such shallow cues.","fun_headline_variants_meta":{"raw":{"variants":["Emotion prompts boost empathy but also verbosity in AI counter-speech","LLM counter-narratives need college-level reading, study says","Smaller AI models can write more accessible hate-speech rebuttals","Hate detectors often misclassify counter-narratives that quote hate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001089,"raw_usage":{"total_tokens":4525,"prompt_tokens":893,"completion_tokens":3632,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":3563}},"tokens_in":509,"tokens_out":3632,"duration_ms":24539,"temperature":1.0,"reasoning_tokens":3563,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:48:35.392253+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human annotation study in which native speakers rate counter-narratives for perceived empathy, warmth, and readability without knowing which prompt produced them would settle whether the NGO-Emotion advantage is real. If human ratings do not show a consistent advantage for the NGO-Emotion condition over NGO-Persona, the paper's central claim would be falsified. Additionally, a simple lexical probe--removing or varying words like 'love', 'happy', 'joy' from the generated texts--would reveal whether the classifier's 'empathetic' labels track those surface cues.","supporting_citations":[{"cited_title":"Human-in-the-Loop for Data Collection: a Multi-Target Counter Narrative Dataset to Fight Online Hate Speech","cited_arxiv_id":"2107.08720","evidence_quote":"Provides the MT-Conan dataset of hate speech and human-authored counter-narratives used as the primary benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the HatEval dataset of real-world hate speech against women and immigrants, used as the secondary, more explicit dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the vanilla prompting approach and an earlier analysis of GPT-generated counter-narratives on MT-Conan, which this paper extends with persona and emotion conditions."}],"review_version":1}