{"id":"e9e3145c-3356-4785-8354-689eee9f9d12","arxiv_id":"2412.07338","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Contextualized counterspeech from LLaMA2-13B using conversation and user history is rated more adequate and persuasive than generic counterspeech, but algorithmic metrics rank configurations inconsistently with humans.","lead":"This paper tests whether AI-generated replies to toxic Reddit comments work better when tailored to the community, the conversation, and the author's history. Crowdsourced raters found that two context-aware configurations beat the generic baseline on adequacy and persuasiveness, while automatic text metrics disagreed with the human rankings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'state-of-the-art generic counterspeech' is never human-evaluated; the only comparison baseline is the unmodified LLaMA2-13B model [Ba], so the headline superiority claim is unsupported for generic SOTA systems.","rationale":"The reader's weakest assumption is the right one. I agree with the CONDITIONAL verdict: the internal human experiment is large and pre-registered, and the specific [Ba Pr Hi] versus [Ba] and [Ba Pr] versus [Ba] comparisons are valid as reported; but the abstract generalizes to 'state-of-the-art generic counterspeech,' which the design never operationalizes. This is load-bearing because the paper's novelty and contribution statement in Sections 1 and 7 rest on beating generic counterspeech, not just beating a base model. The missing baseline is an external-validity gap, not an internal inconsistency; the algorithmic results suggest [Mu] is a strong generic configuration (Table 1), and the poor algorithmic-human correlation means we cannot infer human ratings of [Mu] from indicators. The paper deserves credit for releasing models and pre-registering the experiments, and the within-experiment findings stand on their own. The fix is either adding [Mu] (or another task-tuned generic system) to the human evaluation or rephrasing the claim to refer to the unmodified base model. I would keep the verdict CONDITIONAL: the paper should be accepted only with the claim qualified or the baseline added.","tokens_in":18319,"tokens_out":4880,"duration_ms":48735,"concrete_test":"Run a pre-registered human evaluation with the same protocol and questionnaire as Section 4.2.3, in the non-contextual condition, comparing [Ba], [Mu] (MultiCONAN fine-tuned LLaMA2-13B), [Ba Pr], and [Ba Pr Hi], with sample size following the power analysis in Section 4.2.5. If [Mu] achieves adequacy or persuasiveness scores comparable to or better than [Ba Pr]/[Ba Pr Hi], the abstract's superiority claim over state-of-the-art generic counterspeech is not supported; the claim should be narrowed to the unmodified base model. If [Mu] is clearly worse, the current headline is strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparative claim depends on 'state-of-the-art generic counterspeech' being a meaningful baseline. In the human experiments the only generic baseline is [Ba], defined in Section 4.1 as 'the base LLaMA2-13B model without modifications.' The paper's own fine-tuned configurations [Mu] (MultiCONAN) and [Hs] (RHSI) are described as reproducing state-of-the-art results in automated counterspeech generation (Section 4.1), but neither is included in the seven conditions of the human evaluation (Section 4.2.3, Figures 3-4). The configuration-selection procedure (Section 4.2.2) ranks all 36 configurations algorithmically and selects best/worst per group; for the no-adaptation/no-personalization group, the selected condition is the baseline [Ba] only. [Mu] appears in the human evaluation only inside other configurations (e.g., [Mu Re]), and [Mu Re] is significantly worse than [Ba] on most human-rated aspects. Thus the experiments establish that adding conversation context and comment history to an unmodified instruct model improves perceived adequacy and persuasiveness over that same unmodified model; they do not establish superiority over a generic counterspeech system tuned for the task. The abstract and Section 7 phrase the result as outperforming 'state-of-the-art generic counterspeech,' which overstates what was tested. The Limitations section acknowledges using a single LLM and limited strategies, but it does not flag the absence of a task-tuned generic baseline in the human evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes and evaluates strategies for generating 'contextualized counterspeech' by augmenting an instruction-tuned LLaMA2-13B model with conversation context (previous messages, community fine-tuning) and user personalization (comment history, user summaries). The authors instantiate 36 model configurations, score them with automatic indicators (relevance, diversity, readability, toxicity, adaptation, personalization), select seven configurations for a pre-registered mixed-design crowdsourced human evaluation with two between-subjects conditions (non-contextual and contextual), and compare each configuration against the unmodified LLaMA2-13B baseline [Ba] using nonparametric tests. Results show that [Ba Pr Hi] (previous messages plus comment history) improves adequacy and persuasiveness over [Ba] in the non-contextual condition, and [Ba Pr] improves persuading the author in the contextual condition; algorithmic indicators correlate poorly with human judgments (Kendall tau = -0.05 and -0.43).","tokens_in":18618,"tokens_out":4791,"duration_ms":48407,"significance":"If the central claim is properly scoped, the paper makes a solid empirical contribution: it provides large-scale, pre-registered human evidence that adding conversation context and user history to a base instruct model can improve perceived adequacy and persuasiveness of AI-generated counterspeech, and it documents a stark divergence between algorithmic indicators and human ratings that is important for future evaluation practice. The study's methodological strengths include the large participant pool (about 2,400 per between-subjects condition), a pre-registered protocol, appropriate nonparametric statistics with effect sizes and confidence intervals, and publicly released models and prompts. However, the headline claim that contextualized counterspeech 'significantly outperform[s] state-of-the-art generic counterspeech' is not supported by the human experiments, because the only generic baseline tested is the unmodified [Ba] model; the paper's own task-fine-tuned configurations [Mu] and [Hs], which it describes as reproducing state-of-the-art results, are not human-evaluated standalone.","major_comments":[{"comment":"The central claim that contextualized counterspeech 'can significantly outperform state-of-the-art generic counterspeech in adequacy and persuasiveness' is not supported by the reported human evaluation. In the non-contextual experiment, the only generic baseline is [Ba], defined in Section 4.1 as 'the base LLaMA2-13B model without modifications.' The paper states in Section 4.1 that fine-tuning with MultiCONAN ([Mu]) or RHSI ([Hs]) 'allow[s] us to reproduce state-of-the-art results in automated counterspeech generation,' yet neither [Mu] nor [Hs] is included as a standalone condition in the human evaluation (Section 4.2.3; Figures 3 and 4). The human experiments therefore establish a relative improvement over an unmodified instruct LLM, not over a task-tuned generic counterspeech system. Either the claim should be reworded to specify the actual baseline, or the human evaluation should include [Mu] and [Hs] as generic comparison conditions.","section":"Abstract and Section 6.2.1"},{"comment":"The Limitations paragraph acknowledges that the results rely on a single LLM and a limited set of strategies and that the evaluation is limited by the indicators and crowdsourced judgments, but it does not disclose the absence of a task-tuned generic baseline in the human evaluation. This omission is material because it directly affects the scope of the headline superiority claim. The limitation statement should be extended to note that the human comparison baseline is an unmodified model rather than a state-of-the-art counterspeech system, so readers are not misled about what was tested.","section":"Section 7 (Limitations)"}],"minor_comments":[{"comment":"The 'Adaptation' indicator is defined as 1 - ROUGE between the counterspeech generated by the baseline [Ba] and that generated by each other configuration. This measures divergence from the baseline, not adaptation to the moderation context, and it is used in the configuration-selection super-ranking (Section 4.2.2). The name is misleading; consider renaming it to 'baseline divergence' or providing an explicit justification for why this measure is a valid adaptation proxy.","section":"Section 4.2.1"},{"comment":"The notation for configurations is inconsistent: the text often writes '[Ba Pr Hi]' with spaces, while Table 1 and the figure labels use 'BaPrHi' or 'Ba Pr Hi' variants. Please standardize the notation throughout the manuscript and figure captions.","section":"Table 1 and Figures 3-4"},{"comment":"The significance annotations differ between the two figures (Figure 3 uses only 'p < 0.01', Figure 4 uses three thresholds). Please ensure the caption for Figure 3 also explains any absent symbols and that the thresholds are defined consistently for all panels.","section":"Figure 3 and Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-executed empirical study with a large pre-registered human evaluation, and the negative correlation between algorithmic indicators and human judgments is a valuable secondary result. The main problem is the mismatch between the abstract's claim of outperforming 'state-of-the-art generic counterspeech' and the actual human-evaluated baseline, which is only the unmodified LLaMA2-13B model. This is a load-bearing issue for the paper's central contribution. If the authors either add a human evaluation of the standalone [Mu] and [Hs] generic systems or reword the claim to be accurate, the paper would be a solid contribution. I recommend major revision rather than rejection because the underlying evidence is meaningful and the overstatement is fixable without changing the study's core methodology."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one if you care about how we evaluate LLM-generated moderation content. The core result is real, but the headline claim overreaches slightly.\n\nWhat's new: the systematic combination of community adaptation, conversation context, and user-history/summary personalization for counterspeech, tested across 36 configurations with a genuinely careful human evaluation. Nobody else has put these context dimensions side by side with human judgments; the closest prior work used only message content or a couple of demographic attributes. They also release the models, and the study is pre-registered.\n\nCredit where due: the human evaluation is above the field's usual bar—about 2,400 participants per between-subjects condition, Friedman tests with Wilcoxon/Bonferroni follow-ups, effect sizes with confidence intervals. The result that [Ba Pr Hi] (base model plus conversation context and user history) beats the unmodified LLaMA2-13B baseline on adequacy and persuasiveness is well supported, and [Ba Pr] improves perceived persuasion of the author in the contextual condition. The negative Kendall correlations between algorithmic indicators and human rankings (tau = −0.05, −0.43) are a useful empirical warning to anyone relying on automatic metrics alone.\n\nThe soft spot is in the Abstract and Section 7. 'Outperform state-of-the-art generic counterspeech' is not what the human experiments tested. The only generic baseline was [Ba], the unmodified instruct model. The paper's own fine-tuned [Mu]—which reproduces state-of-the-art automated counterspeech results, per Section 4.1—was never human-evaluated on its own. So the evidence shows that adding conversation and user context to an unmodified instruct model improves perceived adequacy and persuasiveness over that same unmodified model. That's a real finding, but narrower than the headline. The Limitations section doesn't flag this, and it should. Easy fix: add [Mu] as a human-evaluated baseline or soften the claim.\n\nTwo smaller notes. The configuration selection picks algorithmically best/worst per group—but the algorithm anti-correlates with human judgments, so the human-evaluated set is deliberately skewed toward extremes, and the winners ([Ba Pr], [Ba Pr Hi]) were the algorithm's worst picks. The paper surfaces the discrepancy, so it's not hidden, but it does limit how representative the human results are of the full 36-configuration space. And 'adaptation' is defined as 1 − ROUGE against the baseline, which is self-referential by construction; not load-bearing, but worth knowing.\n\nBottom line: send it to review. This is for people working on content moderation, counterspeech generation, and LLM evaluation methodology. The core empirical claim holds, the methodology is solid, and the overclaim is correctable. I'd ask the authors to either strengthen the baseline or tighten the wording, and to discuss the selection-procedure entanglement explicitly.","headline":"A careful, pre-registered human evaluation supports context-aware counterspeech over an unmodified baseline, but the 'state-of-the-art generic counterspeech' headline overclaims—worth sending to review with a request to fix the baseline or the wording.","tokens_in":19210,"tokens_out":6438,"would_cite":true,"duration_ms":57139,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that feeding an LLM contextual information about the conversation and the user who posted a toxic comment yields counterspeech that humans rate as more adequate and persuasive than generic one-size-fits-all AI replies…","keywords":["counterspeech","content moderation","large language models","personalization","adaptation","human evaluation","online toxicity","crowdsourcing"],"falsifier":"Run the same crowdsourced evaluation with the MultiCONAN fine-tuned configuration [Mu] added as an additional generic baseline; if [Mu] matches or beats [Ba Pr Hi] on adequacy and persuasiveness, the contextualization advantage disappears. The paper already computed algorithmic scores for [Mu], so this is a direct extension of its own protocol.","tokens_in":18080,"feed_emoji":"🛡️","tokens_out":4445,"duration_ms":44205,"temperature":0.7,"pith_summary":"The paper tries to establish that counterspeech replies generated by an LLM can be made more effective by feeding them contextual information about the conversation and the person who wrote the toxic message, rather than generating one-size-fits-all replies from the toxic text alone. It reports human evaluations showing that a configuration using conversation history and user comment history outperforms the generic baseline on adequacy and persuasiveness. It also claims that standard automated quality indicators rank counterspeech systems very differently from human raters, so those indicators alone can mislead. If true, this would give platforms a scalable moderation tool that is more persuasive than generic AI replies and would push the field toward evaluation methods that include human judgment.","feed_headline":"Tailored AI counterspeech outranks generic replies in human test","feed_subtitle":"Giving a chatbot conversation history and user comments improved adequacy and persuasiveness over one-size-fits-all replies.","key_machinery":"The machinery is a factorised generation setup: a fixed LLaMA2-13B-Instruct generator modified by seven binary factors ([Ba], [Mu], [Hs], [Re], [Pr], [Hi], [Su]) combined into 36 configurations, with a pre-registered mixed-design crowdsourcing protocol that rates each output on relevance, adequacy, truthfulness, artificiality, and two persuasiveness questions. The factors define what context the model sees; the human experiment isolates the effect of context by showing some raters only the toxic message plus reply and others the same pair with the contextual inputs.","core_discovery":"Across 36 configurations built from an instruction-tuned LLaMA2-13B model, the combination [Ba Pr Hi] — the base model given two preceding conversation messages and ten previous comments by the toxic user — was rated by crowd workers as significantly more adequate and more likely to persuade the toxic author than the unmodified [Ba] baseline in the non-contextual condition; in the contextual condition, [Ba Pr] significantly beat the baseline at persuading the author. Fine-tuned configurations such as [Mu Re], [Hs Hi], [Mu Hs Hi], and [Mu Re Pr Hi] were rated significantly worse than the baseline on most aspects. Rankings from quantitative indicators correlated negatively with human rankings, so the authors conclude that algorithmic metrics and humans assess different qualities of counterspeech.","pith_inferences":["My inference: the finding suggests the bottleneck is not the amount of user data but the model's ability to integrate it, because the best configurations used raw comment history rather than fine-tuned counterspeech knowledge, implying instruction-following may matter more than domain fine-tuning.","My inference: a testable extension is a field-style experiment on Reddit comparing replies from [Ba Pr Hi] against generic replies for actual changes in author behaviour, such as editing, deleting, or replying more civilly, which crowdsourcing ratings cannot capture.","My inference: the negative correlation between algorithmic and human rankings implies that any automated leaderboard for counterspeech may rank systems opposite to human preference, so combining both evaluation modes is a validity requirement rather than optional polish."],"forward_implications":["Conversation history and user comment history are usable inputs for more persuasive AI-generated counterspeech, with no measured loss on relevance, truthfulness, or civility.","Automated indicators such as ROUGE, readability, toxicity, and style similarity should not be used alone to rank counterspeech systems, since their ranking disagreed with human judgments in this study.","Showing human evaluators the context behind a counterspeech reply changes their ratings and compresses quality differences between configurations, meaning context matters for how interventions are perceived.","Configurations that combine many factors tend to drift from instructions and produce inadequate or meaningless replies, so simpler, targeted context may be safer than maximal context.","Future use of larger models, which handle multiple instructions better, is likely to improve contextualized counterspeech further."],"supporting_citations":[{"why":"Supplies the base instruction-tuned LLaMA2-13B model used as the counterspeech generator and as the baseline configuration.","marker":"[70]"},{"why":"Provides the MultiCONAN hate speech–counterspeech pairs used for the [Mu] fine-tuning factor.","marker":"[23]"},{"why":"Provides the Reddit hate-speech intervention dataset used for the [Hs] fine-tuning factor.","marker":"[56]"},{"why":"Supplies the Perspective API used to measure toxicity in the algorithmic evaluation.","marker":"[49]"},{"why":"Provides the prior human-evaluation approach for persuasiveness of counterspeech that the two-question persuasiveness measure follows.","marker":"[41]"},{"why":"Motivates personalized and adapted moderation interventions, the premise for generating contextualized counterspeech.","marker":"[18]"},{"why":"Represents the prior attempt at contextualized counterspeech that this paper claims to surpass with the first successful human-evaluated results.","marker":"[8]"},{"why":"Supplies the survey-based framing of counterspeech properties and evaluation used to define adequacy, truthfulness, and related attributes.","marker":"[5]"},{"why":"Establishes the comparative state of the art for generic LLM counterspeech generation that the baseline configuration is meant to represent.","marker":"[67]"}],"fun_headline_variants":["Context-aware AI replies beat generic counterspeech","Personalized AI counterspeech wins human ratings","AI counterspeech with context persuades better than generic","Tailored AI replies with user context beat generic ones"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on treating a plain, unmodified chatbot as the state-of-the-art generic counterspeech; a stronger generic baseline was never tested with human raters, so the advantage of context could shrink if one were added.","fun_headline_variants_meta":{"raw":{"variants":["Context-aware AI replies beat generic counterspeech","Personalized AI counterspeech wins human ratings","AI counterspeech with context persuades better than generic","Tailored AI replies with user context beat generic ones"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000724,"raw_usage":{"total_tokens":3222,"prompt_tokens":894,"completion_tokens":2328,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":2269}},"tokens_in":510,"tokens_out":2328,"duration_ms":23410,"temperature":1.0,"reasoning_tokens":2269,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:55:20.484367+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same crowdsourced evaluation with the MultiCONAN fine-tuned configuration [Mu] added as an additional generic baseline; if [Mu] matches or beats [Ba Pr Hi] on adequacy and persuasiveness, the contextualization advantage disappears. The paper already computed algorithmic scores for [Mu], so this is a direct extension of its own protocol.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the base instruction-tuned LLaMA2-13B model used as the counterspeech generator and as the baseline configuration."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the MultiCONAN hate speech–counterspeech pairs used for the [Mu] fine-tuning factor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Reddit hate-speech intervention dataset used for the [Hs] fine-tuning factor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Perspective API used to measure toxicity in the algorithmic evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior human-evaluation approach for persuasiveness of counterspeech that the two-question persuasiveness measure follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates personalized and adapted moderation interventions, the premise for generating contextualized counterspeech."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the comparative state of the art for generic LLM counterspeech generation that the baseline configuration is meant to represent."}],"review_version":1}