Pith. sign in

REVIEW 2 major objections 3 minor 87 references

Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation

T0 review · 2 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper claims that feeding an LLM contextual information about the conversation and the user who posted a toxic comment yields counterspeech that humans rate as more adequate and persuasive than generic one-size-fits-all AI replies…

desk verdict A careful, pre-registered human evaluation supports context-aware counterspeech over an unmodified baseline, but the 'state-of-the-art generic counterspeech' headline overclaims—worth sending to review with a request to fix the baseline or the wording. read the letter →

arxiv 2412.07338 v3 pith:NOC23TP5 submitted 2024-12-10 cs.HC cs.AIcs.SI

classification cs.HCcs.AIcs.SI
keywords counterspeechcontentmoderationlargelanguagemodelspersonalizationadaptationhumanevaluationonlinetoxicitycrowdsourcing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that counterspeech replies generated by an LLM can be made more effective by feeding them contextual information about the conversation and the person who wrote the toxic message, rather than generating one-size-fits-all replies from the toxic text alone. It reports human evaluations showing that a configuration using conversation history and user comment history outperforms the generic baseline on adequacy and persuasiveness. It also claims that standard automated quality indicators rank counterspeech systems very differently from human raters, so those indicators alone can mislead. If true, this would give platforms a scalable moderation tool that is more persuasive than generic AI replies and would push the field toward evaluation methods that include human judgment.

What carries the argument

The machinery is a factorised generation setup: a fixed LLaMA2-13B-Instruct generator modified by seven binary factors ([Ba], [Mu], [Hs], [Re], [Pr], [Hi], [Su]) combined into 36 configurations, with a pre-registered mixed-design crowdsourcing protocol that rates each output on relevance, adequacy, truthfulness, artificiality, and two persuasiveness questions. The factors define what context the model sees; the human experiment isolates the effect of context by showing some raters only the toxic message plus reply and others the same pair with the contextual inputs.

What would settle it

Run the same crowdsourced evaluation with the MultiCONAN fine-tuned configuration [Mu] added as an additional generic baseline; if [Mu] matches or beats [Ba Pr Hi] on adequacy and persuasiveness, the contextualization advantage disappears. The paper already computed algorithmic scores for [Mu], so this is a direct extension of its own protocol.

Watch

Extended reading notes

Core claim

Across 36 configurations built from an instruction-tuned LLaMA2-13B model, the combination [Ba Pr Hi] — the base model given two preceding conversation messages and ten previous comments by the toxic user — was rated by crowd workers as significantly more adequate and more likely to persuade the toxic author than the unmodified [Ba] baseline in the non-contextual condition; in the contextual condition, [Ba Pr] significantly beat the baseline at persuading the author. Fine-tuned configurations such as [Mu Re], [Hs Hi], [Mu Hs Hi], and [Mu Re Pr Hi] were rated significantly worse than the baseline on most aspects. Rankings from quantitative indicators correlated negatively with human rankings, so the authors conclude that algorithmic metrics and humans assess different qualities of counterspeech.

Load-bearing premise

The claim rests on treating a plain, unmodified chatbot as the state-of-the-art generic counterspeech; a stronger generic baseline was never tested with human raters, so the advantage of context could shrink if one were added.

Editorial extensions

If this is right

  • Conversation history and user comment history are usable inputs for more persuasive AI-generated counterspeech, with no measured loss on relevance, truthfulness, or civility.
  • Automated indicators such as ROUGE, readability, toxicity, and style similarity should not be used alone to rank counterspeech systems, since their ranking disagreed with human judgments in this study.
  • Showing human evaluators the context behind a counterspeech reply changes their ratings and compresses quality differences between configurations, meaning context matters for how interventions are perceived.
  • Configurations that combine many factors tend to drift from instructions and produce inadequate or meaningless replies, so simpler, targeted context may be safer than maximal context.
  • Future use of larger models, which handle multiple instructions better, is likely to improve contextualized counterspeech further.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the finding suggests the bottleneck is not the amount of user data but the model's ability to integrate it, because the best configurations used raw comment history rather than fine-tuned counterspeech knowledge, implying instruction-following may matter more than domain fine-tuning.
  • My inference: a testable extension is a field-style experiment on Reddit comparing replies from [Ba Pr Hi] against generic replies for actual changes in author behaviour, such as editing, deleting, or replying more civilly, which crowdsourcing ratings cannot capture.
  • My inference: the negative correlation between algorithmic and human rankings implies that any automated leaderboard for counterspeech may rank systems opposite to human preference, so combining both evaluation modes is a validity requirement rather than optional polish.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper proposes and evaluates strategies for generating 'contextualized counterspeech' by augmenting an instruction-tuned LLaMA2-13B model with conversation context (previous messages, community fine-tuning) and user personalization (comment history, user summaries). The authors instantiate 36 model configurations, score them with automatic indicators (relevance, diversity, readability, toxicity, adaptation, personalization), select seven configurations for a pre-registered mixed-design crowdsourced human evaluation with two between-subjects conditions (non-contextual and contextual), and compare each configuration against the unmodified LLaMA2-13B baseline [Ba] using nonparametric tests. Results show that [Ba Pr Hi] (previous messages plus comment history) improves adequacy and persuasiveness over [Ba] in the non-contextual condition, and [Ba Pr] improves persuading the author in the contextual condition; algorithmic indicators correlate poorly with human judgments (Kendall tau = -0.05 and -0.43).

Significance. If the central claim is properly scoped, the paper makes a solid empirical contribution: it provides large-scale, pre-registered human evidence that adding conversation context and user history to a base instruct model can improve perceived adequacy and persuasiveness of AI-generated counterspeech, and it documents a stark divergence between algorithmic indicators and human ratings that is important for future evaluation practice. The study's methodological strengths include the large participant pool (about 2,400 per between-subjects condition), a pre-registered protocol, appropriate nonparametric statistics with effect sizes and confidence intervals, and publicly released models and prompts. However, the headline claim that contextualized counterspeech 'significantly outperform[s] state-of-the-art generic counterspeech' is not supported by the human experiments, because the only generic baseline tested is the unmodified [Ba] model; the paper's own task-fine-tuned configurations [Mu] and [Hs], which it describes as reproducing state-of-the-art results, are not human-evaluated standalone.

major comments (2)
  1. [Abstract and Section 6.2.1] The central claim that contextualized counterspeech 'can significantly outperform state-of-the-art generic counterspeech in adequacy and persuasiveness' is not supported by the reported human evaluation. In the non-contextual experiment, the only generic baseline is [Ba], defined in Section 4.1 as 'the base LLaMA2-13B model without modifications.' The paper states in Section 4.1 that fine-tuning with MultiCONAN ([Mu]) or RHSI ([Hs]) 'allow[s] us to reproduce state-of-the-art results in automated counterspeech generation,' yet neither [Mu] nor [Hs] is included as a standalone condition in the human evaluation (Section 4.2.3; Figures 3 and 4). The human experiments therefore establish a relative improvement over an unmodified instruct LLM, not over a task-tuned generic counterspeech system. Either the claim should be reworded to specify the actual baseline, or the human evaluation should include [Mu] and [Hs] as generic comparison conditions.
  2. [Section 7 (Limitations)] The Limitations paragraph acknowledges that the results rely on a single LLM and a limited set of strategies and that the evaluation is limited by the indicators and crowdsourced judgments, but it does not disclose the absence of a task-tuned generic baseline in the human evaluation. This omission is material because it directly affects the scope of the headline superiority claim. The limitation statement should be extended to note that the human comparison baseline is an unmodified model rather than a state-of-the-art counterspeech system, so readers are not misled about what was tested.
minor comments (3)
  1. [Section 4.2.1] The 'Adaptation' indicator is defined as 1 - ROUGE between the counterspeech generated by the baseline [Ba] and that generated by each other configuration. This measures divergence from the baseline, not adaptation to the moderation context, and it is used in the configuration-selection super-ranking (Section 4.2.2). The name is misleading; consider renaming it to 'baseline divergence' or providing an explicit justification for why this measure is a valid adaptation proxy.
  2. [Table 1 and Figures 3-4] The notation for configurations is inconsistent: the text often writes '[Ba Pr Hi]' with spaces, while Table 1 and the figure labels use 'BaPrHi' or 'Ba Pr Hi' variants. Please standardize the notation throughout the manuscript and figure captions.
  3. [Figure 3 and Figure 4] The significance annotations differ between the two figures (Figure 3 uses only 'p < 0.01', Figure 4 uses three thresholds). Please ensure the caption for Figure 3 also explains any absent symbols and that the thresholds are defined consistently for all panels.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central result rests on pre-registered human evaluations and does not reduce to its inputs by construction.

full rationale

The paper's main claim—that contextualized counterspeech can outperform generic counterspeech on adequacy and persuasiveness—is supported by direct human ratings from a pre-registered mixed-design crowdsourcing experiment, not by a derivation from fitted parameters or self-citations. The comparisons in Figures 3 and 4 are against the [Ba] configuration, an unmodified LLaMA2-13B-Instruct model; the wording 'state-of-the-art generic counterspeech' in the abstract may overstate the strength of this baseline, since no task-tuned generic counterspeech system was human-evaluated, but this is a validity or framing concern rather than a circular one. The 'adaptation' indicator in Section 4.2.1 is defined as 1 minus ROUGE between baseline and configuration outputs, so it is tautologically related to output divergence, but the paper does not rely on this indicator to establish persuasiveness; it explicitly reports that algorithmic and human rankings diverge (Kendall tau near zero or negative in Figure 6). No fitted parameter is renamed as a prediction, no load-bearing uniqueness theorem is imported from the authors' prior work, and the authors' self-citations appear only as background or related work. The derivation chain is therefore self-contained with respect to the empirical evaluation, and no circular step is exhibited.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The paper's claim rests on an empirical pipeline: selecting 128 toxic comments with a toxicity threshold, generating replies using hand-chosen context sizes, and evaluating 20 representative messages per configuration via crowdsourcing. These choices are not fitted to the outcome, but they shape the result, so they are listed as free parameters. The key domain assumptions are that Likert-scale persuasiveness ratings approximate real persuasive impact, and that Reddit political subreddits stand in for online toxicity more broadly. No invented theoretical entities are introduced.

free parameters (6)
  • toxicity threshold for selecting toxic comments = 0.5 Perspective API score
    Used in Section 5 to select the 128 toxic comments for counterspeech generation; changing the threshold would alter the evaluation dataset and potentially the results.
  • number of previous messages for conversation context [Pr] = 2 parent messages
    Section 4.1.1: up to two parent messages m_i-1 and m_i-2 are prepended; the amount of conversation context is hand-chosen.
  • number of user comments for comment history [Hi] = 10
    Section 4.1.2: ten previous messages from the author are prepended to personalize the generator.
  • number of user comments for user summary [Su] = 20
    Section 4.1.2: twenty comments are summarized into a user profile that is then provided to the generator.
  • representative messages selected for human evaluation = 20 per configuration
    Section 4.2.2: only 20 of 128 generated messages per configuration were rated by humans; selection by centroid proximity may not capture the full output distribution.
  • community adaptation fine-tuning sample size [Re] = about 7,500 comment-reply pairs
    Section 5: size of the stratified sample used to adapt the model to Reddit political conversational style.
assumptions (5)
  • domain assumption Perceived persuasiveness measured by 5-point Likert ratings on Amazon Mechanical Turk approximates real-world persuasive effectiveness of counterspeech.
    The central claim about persuasiveness is based on ratings of likelihood to persuade, not on observed behavior change; this assumption enters in Sections 4.2.3 and 6.2.
  • domain assumption Reddit political subreddits and the selected 128 toxic comments represent the broader online toxicity context for counterspeech.
    Section 5 limits the dataset to five US-politics subreddits; results may not generalize to other platforms, languages, or topic domains.
  • domain assumption Google Perspective API toxicity score is a valid proxy for toxicity in comment selection.
    Section 5 uses comments with toxicity at least 0.5 as the toxic set; the API is a model with known imperfections.
  • standard math Non-parametric statistical tests (Friedman, Wilcoxon, Mann-Whitney) with Bonferroni correction are appropriate for the mixed-design experiment.
    Section 4.2.4 specifies these tests for within- and between-subjects comparisons; they do not rely on normality assumptions.
  • domain assumption LLaMA2-13B instruction-tuned models consistently generate human-quality text, so basic fluency and grammaticality do not need separate evaluation.
    Section 3 explicitly sets aside fluency and grammaticality based on modern LLM capabilities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation." pith.science (2026). https://pith.science/paper/NOC23TP5

@misc{pith2026241207338,
  author       = {Pith},
  title        = {Pith review of: Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NOC23TP5}},
  note         = {Machine review of arXiv:2412.07338}
}
read the original abstract

AI-generated counterspeech offers a promising and scalable strategy to curb online toxicity through direct replies that promote civil discourse. However, current counterspeech is one-size-fits-all, lacking adaptation to the moderation context and the users involved. We propose and evaluate multiple strategies for generating tailored counterspeech that is adapted to the moderation context and personalized for the moderated user. We instruct an LLaMA2-13B model to generate counterspeech, experimenting with various configurations based on different contextual information and fine-tuning strategies. We identify the configurations that generate persuasive counterspeech through a combination of quantitative indicators and human evaluations collected via a pre-registered mixed-design crowdsourcing experiment. Results show that contextualized counterspeech can significantly outperform state-of-the-art generic counterspeech in adequacy and persuasiveness, without compromising other characteristics. Our findings also reveal a poor correlation between quantitative indicators and human evaluations, suggesting that these methods assess different aspects and highlighting the need for nuanced evaluation methodologies. The effectiveness of contextualized AI-generated counterspeech and the divergence between human and algorithmic evaluations underscore the importance of increased human-AI collaboration in content moderation.

Figures

Figures reproduced from arXiv: 2412.07338 by the authors.

Figure 1
Figure 1. Current AI-generated counterspeech only leverages [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Algorithmic evaluation results for each factor. For each factor ( [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Human evaluation results (non-contextual condition). Effect sizes and confidence intervals of the scores assigned to several configurations compared to the baseline. Statistical significance: ***: 𝑝 < 0.01. BaPr MuRe HsHi MuHsHi BaPrHi MuRePrHi [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Human evaluation results (contextual condition). Effect sizes and confidence intervals of the scores assigned to several configurations compared to the baseline. Statistical significance: ***: 𝑝 < 0.01, **: 𝑝 < 0.05, *: 𝑝 < 0.1 [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Differences in human evaluation results between the [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Aggregated rankings of the selected configurations, [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Human evaluation results (non-contextual condition) based on answers from those participants who reported using social media “very often”. Statistical significance: ***: 𝑝 < 0.01, **: 𝑝 < 0.05. • Personalization. While adaptation focuses on the broad modera￾tion contex…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

87 extracted references · 67 canonical work pages

  1. [1]

    Ana Aleksandric, Sayak Saha Roy, Hanani Pankaj, Gabriela Mustata Wilson, and Shirin Nilizadeh. 2024. Users’ behavioral and emotional response to toxicity in Twitter conversations. In AAAI ICWSM

  2. [2]

    Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. The pushshift Reddit dataset. In AAAI ICWSM

  3. [3]

    Tilman Beck, Hendrik Schuff, Anne Lauscher, and Iryna Gurevych. 2024. Sensi- tivity, performance, robustness: Deconstructing the effect of sociodemographic prompting. In EACL

  4. [4]

    Michał Bilewicz, Patrycja Tempska, Gniewosz Leliwa, Maria Dowgiałło, Michalina Tańska, Rafał Urbaniak, and Michał Wroczyński. 2021. Artificial intelligence against hate: Intervention reducing verbal aggression in the social network environment. Aggressive Behavior 47, 3 (2021)

  5. [5]

    Helena Bonaldi, Yi-Ling Chung, Gavin Abercrombie, and Marco Guerini. 2024. NLP for counterspeech against hate: A survey and how-to guide. In NAACL

  6. [6]

    Simon Martin Breum, Daniel Vædele Egdal, Victor Gram Mortensen, Anders Gio- vanni Møller, and Luca Maria Aiello. 2024. The persuasive power of large language models. In AAAI ICWSM

  7. [7]

    Dominique Brunato, Andrea Cimino, Felice Dell’Orletta, Giulia Venturi, and Simonetta Montemagni. 2020. Profiling-UD: A tool for linguistic profiling of texts. In LREC

  8. [8]

    Dominik Bär, Abdurahman Maarouf, and Stefan Feuerriegel. 2024. Generative AI may backfire for counterspeech. arXiv:2411.14986 (2024)

Show all 87 references
  1. [9]

    Eshwar Chandrasekharan, Shagun Jhaver, Amy Bruckman, and Eric Gilbert

  2. [10]

    Hyundong Cho, Shuai Liu, Taiwei Shi, Darpan Jain, Basem Rizk, Yuyang Huang, Zixun Lu, Nuan Wen, Jonathan Gratch, Emilio Ferrara, and Jonathan May. 2024. Can language model moderators improve the health of online discourse? NAACL (2024)

  3. [11]

    Yi-Ling Chung, Gavin Abercrombie, Florence Enock, Jonathan Bright, and Ver- ena Rieser. 2023. Understanding counterspeech for online harm mitigation. arXiv:2307.04761 (2023)

  4. [12]

    Yi-Ling Chung, Elizaveta Kuzmenko, Serra Sinem Tekiroglu, and Marco Guerini

  5. [13]

    Yi-Ling Chung, Serra Sinem Tekiroğlu, and Marco Guerini. 2021. Towards knowledge-grounded counter narrative generation for hate speech. In ACL- IJCNLP

  6. [14]

    Lorenzo Cima, Benedetta Tessa, Stefano Cresci, Amaury Trujillo, and Marco Avvenuti. 2024. Investigating the heterogenous effects of a massive content moderation intervention via Difference-in-Differences. arXiv:2411.04037 (2024)

  7. [15]

    Lorenzo Cima, Amaury Trujillo, Marco Avvenuti, and Stefano Cresci. 2024. The Great Ban: Efficacy and unintended consequences of a massive deplatforming operation on Reddit. In ACM WebSci Companion

  8. [16]

    William Jay Conover. 1999. Practical nonparametric statistics. John Wiley & Sons

  9. [17]

    Thomas H Costello, Gordon Pennycook, and David G Rand. 2024. Durably reducing conspiracy beliefs through dialogues with AI. Science 385 (2024)

  10. [18]

    Stefano Cresci, Amaury Trujillo, and Tiziano Fagni. 2022. Personalized interven- tions for online moderation. In ACM Hypertext

  11. [19]

    I’m a Professor, which isn’t usually a dangerous job

    Periwinkle Doerfler, Andrea Forte, Emiliano De Cristofaro, Gianluca Stringhini, Jeremy Blackburn, and Damon McCoy. 2021. “I’m a Professor, which isn’t usually a dangerous job”: Internet-facilitated harassment and its impact on researchers. In ACM CSCW

  12. [20]

    Mekselina Doğanç and Ilia Markov. 2023. From generic to personalized: Investi- gating strategies for generating targeted counter narratives against hate speech. In ACL CS4OA

  13. [21]

    Cynthia Dwork, Chris Hays, Jon Kleinberg, and Manish Raghavan. 2024. Content moderation and the formation of online communities: A theoretical framework. In The ACM Web Conf

  14. [22]

    European Commission. 2019. Progress on combating hate speech online through the EU Code of Conduct . https://data.consilium.europa.eu/doc/document/ST- 12522-2019-INIT/en/pdf

  15. [23]

    Margherita Fanton, Helena Bonaldi, Serra Sinem Tekiroğlu, and Marco Guerini

  16. [24]

    Kazuaki Furumai, Roberto Legaspi, Julio Vizcarra, Yudai Yamazaki, Yasutaka Nishimura, Sina J Semnani, Kazushi Ikeda, Weiyan Shi, and Monica S Lam. 2024. Zero-shot persuasive chatbots with LLM-generated strategies and information retrieval. EMNLP (2024)

  17. [25]

    John D Gallacher, Marc W Heerdink, and Miles Hewstone. 2021. Online engage- ment between opposing political protest groups via social media is linked to physical violence of offline encounters. Social Media + Society 7, 1 (2021)

  18. [26]

    Joshua Garland, Keyan Ghazi-Zahedi, Jean-Gabriel Young, Laurent Hébert- Dufresne, and Mirta Galesic. 2022. Impact and dynamics of hate and counter speech online. EPJ Data Science 11, 1 (2022)

  19. [27]

    Panagiotis Germanakos, Marios Belk, et al. 2016. Human-centred web adaptation and personalization. Springer

  20. [28]

    Tarleton Gillespie. 2020. Content moderation, AI, and the question of scale. Big Data & Society 7, 2 (2020)

  21. [29]

    Tommaso Giorgi, Lorenzo Cima, Tiziano Fagni, Marco Avvenuti, and Stefano Cresci. 2025. Human and LLM biases in hate speech annotations: A socio- demographic analysis of annotators and targets. AAAI ICWSM (2025)

  22. [30]

    Natasha Goel, Thomas Bergeron, Blake Lee-Whiting, Thomas Galipeau, Danielle Bohonos, Sarah Lachance, Sonja Savolainen, Clareta Treger, and Eric Merkley

  23. [31]

    Pierpaolo Goffredo, Valerio Basile, Bianca Cepollaro, Viviana Patti, et al. 2022. Counter-TWIT: An Italian corpus for online counterspeech in ecological contexts. In ACL WOAH

  24. [32]

    Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Stamos, and Michael Tomz

  25. [33]

    Jarod Govers, Eduardo Velloso, Vassilis Kostakos, and Jorge Goncalves. 2024. AI-Driven Mediation Strategies for Audience Depolarisation in Online Debates. In ACM CHI

  26. [34]

    Kobi Hackenburg and Helen Margetts. 2024. Evaluating the persuasive influence of political microtargeting with large language models.Proceedings of the National Academy of Sciences 121, 24 (2024)

  27. [35]

    Kobi Hackenburg, Ben M Tappin, Paul Röttger, Scott Hale, Jonathan Bright, and Helen Margetts. 2024. Evidence of a log scaling law for political persuasion with large language models. arXiv:2406.14508 (2024)

  28. [36]

    Sadaf MD Halim, Saquib Irtiza, Yibo Hu, Latifur Khan, and Bhavani Thuraising- ham. 2023. WokeGPT: Improving counterspeech generation against online hate speech by intelligently augmenting datasets using a novel metric. InIEEE IJCNN

  29. [37]

    How persuasive is AI-generated propaganda? PNAS Nexus 3, 2 (2024)

  30. [38]

    Sabit Hassan and Malihe Alikhani. 2023. DisCGen: A framework for discourse- informed counterspeech generation. In IJCNLP-AACL

  31. [39]

    Bing He, Mustaque Ahamad, and Srijan Kumar. 2023. Reinforcement learning- based counter-misinformation response generation: A case study of COVID-19 vaccine misinformation. In The ACM Web Conf

  32. [40]

    Amey Hengle, Aswini Kumar, Anil Bandhakavi, and Tanmoy Chakraborty. 2025. CSEval: Towards automated, multi-dimensional, and reference-free counter- speech evaluation using auto-calibrated LLMs. arXiv:2501.17581 (2025)

  33. [41]

    Lingzi Hong, Pengcheng Luo, Eduardo Blanco, and Xiaoying Song. 2024. Outcome- constrained large language models for countering hate speech. EMNLP (2024)

  34. [42]

    Dominik Hangartner, Gloria Gennaro, Sary Alasiri, Nicholas Bahrich, Alexandra Bornhoft, Joseph Boucher, Buket Buse Demirci, Laurenz Derksen, Aldo Hall, Matthias Jochum, et al. 2021. Empathy-based counterspeech can reduce racist hate speech in a social media field experiment. P...

  35. [43]

    Evey Jiaxin Huang, Abhraneel Sarma, Sohyeon Hwang, Eshwar Chandrasekha- ran, and Stevie Chancellor. 2024. Opportunities, tensions, and challenges in computational approaches to addressing online harassment. In ACM DIS

  36. [44]

    Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara

  37. [45]

    Shuyu Jiang, Wenyi Tang, Xingshu Chen, Rui Tang, Haizhou Wang, and Wenxian Wang. 2025. ReZG: Retrieval-augmented zero-shot counter narrative generation for hate speech. Neurocomputing 620 (2025), 129140

  38. [46]

    JP Kincaid. 1975. Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel. Chief of Naval Technical Training (1975)

  39. [47]

    Manoel Horta Ribeiro, Shagun Jhaver, Savvas Zannettou, Jeremy Blackburn, Gianluca Stringhini, Emiliano De Cristofaro, and Robert West. 2021. Do platform migrations compromise content moderation? Evidence from r/The_Donald and r/Incels. In ACM CSCW

  40. [48]

    Rohan Leekha, Olga Simek, and Charlie Dagli. 2024. War of Words: Harnessing the Potential of Large Language Models and Retrieval Augmented Generation to Classify, Counter and Diffuse Hate Speech. In AAAI FLAIRS

  41. [49]

    Alyssa Lees, Vinh Q Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman. 2022. A new generation of perspective API: Efficient multilingual character-level transformers. In ACM KDD

  42. [50]

    NAACL (2024)

    PersonaLLM: Investigating the ability of large language models to express personality traits. NAACL (2024)

  43. [51]

    Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out

  44. [52]

    Kevin Munger. 2017. Tweetment effects on the tweeted: Experimentally reducing racist harassment. Political Behavior 39 (2017)

  45. [53]

    2020.Statistical reasoning in the behavioral sciences

    Bruce M King, Patrick J Rosopa, and Edward W Minium. 2020.Statistical reasoning in the behavioral sciences . John Wiley & Sons

  46. [54]

    Dino Pedreschi, Luca Pappalardo, Emanuele Ferragina, Ricardo Baeza-Yates, Albert-László Barabási, Frank Dignum, Virginia Dignum, Tina Eliassi-Rad, Fosca Giannotti, János Kertész, et al. 2024. Human-AI coevolution. Artificial Intelligence (2024)

  47. [55]

    Vasyl Pihur, Susmita Datta, and Somnath Datta. 2009. RankAggreg, an R package for weighted rank aggregation. BMC Bioinformatics 10 (2009)

  48. [56]

    Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2024. Pre-trained language models for text generation: A survey. ACM Computing Surveys 56, 9 (2024)

  49. [57]

    Ganesh Prasath Ramani, Shirish Karande, Yash Bhatia, et al. 2024. Persuasion games using large language models. arXiv:2408.15879 (2024)

  50. [58]

    Sarah T Roberts. 2019. Behind the screen: Content moderation in the shadows of social media. Yale University Press

  51. [59]

    Paola Pascual-Ferrá, Neil Alperstein, Daniel J Barnett, and Rajiv N Rimal. 2021. Toxicity and verbal aggression on social media: Polarized discourse on wearing face masks during the COVID-19 pandemic. Big Data & Society 8, 1 (2021). WWW ’25, April 28-May 2, 2025, Sydney, NSW, ...

  52. [60]

    Punyajoy Saha, Kanishk Singh, Adarsh Kumar, Binny Mathew, and Animesh Mukherjee. 2022. CounterGeDi: A controllable approach to generate polite, detoxified and emotional counterspeech. In IJCAI

  53. [61]

    Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. 2024. On the conversational persuasiveness of large language models: A randomized controlled trial. arXiv:2403.14380 (2024)

  54. [62]

    Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding, and William Yang Wang

  55. [63]

    In EMNLP-IJCNLP

    A benchmark dataset for learning to intervene in online hate speech. In EMNLP-IJCNLP

  56. [64]

    Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The psychological well-being of content moderators: The emotional labor of commercial moderation and avenues for improving support. In ACM CHI

  57. [65]

    Derald Wing Sue, Sarah Alsaidi, Michael N Awad, Elizabeth Glaeser, Cassandra Z Calle, and Narolyn Mendez. 2019. Disarming racial microaggressions: Microinter- vention strategies for targets, White allies, and bystanders.American Psychologist 74, 1 (2019)

  58. [66]

    Koustuv Saha, Eshwar Chandrasekharan, and Munmun De Choudhury. 2019. Prevalence and psychological effects of hateful speech in online college commu- nities. In ACM WebSci

  59. [67]

    Serra Sinem Tekiroglu, Helena Bonaldi, Margherita Fanton, and Marco Guerini

  60. [68]

    Serra Sinem Tekiroğlu, Yi-Ling Chung, and Marco Guerini. 2020. Generating counter narratives against online hate speech: Data and strategies. In ACL

  61. [69]

    Carla Schieb and Mike Preuss. 2016. Governing hate speech by means of coun- terspeech on Facebook. In ICA

  62. [70]

    Weiyan Shi, Xuewei Wang, Yoo Jung Oh, Jingwen Zhang, Saurav Sahay, and Zhou Yu. 2020. Effects of persuasive dialogues: Testing bot identities and inquiry strategies. In ACM CHI

  63. [71]

    Amaury Trujillo and Stefano Cresci. 2022. Make Reddit Great Again: Assessing community effects of moderation interventions on r/The_Donald. InACM CSCW

  64. [72]

    Amaury Trujillo and Stefano Cresci. 2023. One of many: Assessing user-level effects of moderation interventions on r/The_Donald. In ACM WebSci

  65. [73]

    Madiha Tabassum, Alana Mackey, Ashley Schuett, and Ada Lerner. 2024. Investi- gating moderation challenges to combating hate and harassment: The case of Mod-Admin power dynamics and feature misuse on Reddit. In USENIX

  66. [74]

    Siyi Wang, Qi Deng, Shiwei Feng, Hong Zhang, and Chao Liang. 2024. A survey on rank aggregation. In IJCAI

  67. [75]

    Using pre-trained language models for producing counter narratives against hate speech: A comparative study. In ACL

  68. [76]

    Xinchen Yu, Eduardo Blanco, and Lingzi Hong. 2024. Hate cannot drive out hate: Forecasting conversation incivility following replies to hate speech. In AAAI ICWSM

  69. [77]

    Benedetta Tessa, Lorenzo Cima, Amaury Trujillo, Marco Avvenuti, and Stefano Cresci. 2024. Beyond trial-and-error: Predicting user abandonment after a mod- eration intervention. arXiv:2404.14846 (2024)

  70. [78]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al . 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288 (2023)

  71. [79]

    very often

    Aneta Zugecova, Dominik Macko, Ivan Srba, Robert Moro, Jakub Kopal, Katarina Marcincinova, and Matus Mesarcik. 2024. Evaluation of LLM vulnerabilities to being misused for personalized disinformation generation. arXiv:2412.13666 (2024). A Summary of relevant Works Table 2 repo...

  72. [81]

    Amaury Trujillo, Tiziano Fagni, and Stefano Cresci. 2025. The DSA Transparency Database: Auditing self-reported moderation actions by social media. In ACM CSCW

  73. [83]

    Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017. Ex machina: Personal attacks seen at scale. In The ACM Web Conf

  74. [85]

    Wanzheng Zhu and Suma Bhat. 2021. Generate, Prune, Select: A pipeline for counterspeech generation against online hate speech. In ACL-IJCNLP

  75. [86]

    Irune Zubiaga, Aitor Soroa, and Rodrigo Agerri. 2024. A LLM-based ranking method for the evaluation of automatic counter-narrative generation. InEMNLP

  76. [2019]

    CONAN – COunter NArratives through Nichesourcing: A multilingual dataset of responses to fight online hate speech. In ACL

  77. [2021]

    In ACL-IJCNLP

    Human-in-the-Loop for data collection: A multi-target counter narrative dataset to fight online hate speech. In ACL-IJCNLP

  78. [2022]

    ACM TOCHI 29, 4 (2022)

    Quarantined! Examining the effects of a community-wide moderation intervention on Reddit. ACM TOCHI 29, 4 (2022)

  79. [2024]

    Artificial influence? Comparing AI and human persuasion in reducing belief certainty. (2024). https://doi.org/10.31219/osf.io/2vh4k

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.