{"id":"e15e9b3c-de70-4852-acd9-f456f1e9c6b0","arxiv_id":"2607.14491","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In seven state-attributed influence operations, only ~19% of detector-flagged hostile content is identity-directed and dehumanizing/inciting hate; the rest is partisan or geopolitical invective, and the mix varies by campaign.","lead":"A new analysis of 25 million tweets from seven state-backed influence campaigns argues that standard 'hate' rates overstate hate roughly twofold, because the detectors actually measure a broader mix of partisan and geopolitical invective. The paper decomposes hostile content into three construct types and shows the mix varies by operation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing input is the LLM-derived target-group attribute driving 69% of typing assignments; it is unvalidated at scale, so the 50.1% identity share and the 2x overstatement could shift.","rationale":"The reader's weakest assumption is the right one. The central empirical claim—that a broad-validated hostility gate is not a hate detector and overstates hate by about 2x on these operations—rests on a decomposition of the 5,457 gate positives into identity/partisan/geopolitical. That decomposition is a deterministic rule over an 11-dimension LLM characterization, and the target-group dimension is decisive for 69.0% of assignments. The paper merits credit for validating the broad gate against human gold (kappa=0.82, precision 0.96 on the 100-item set), for validating the typing rule against an expert at kappa=0.52 and showing it lies within the human-human agreement band on contested items, for the two-prompt conservative census, for the external IRA role corpus, and for the candid limitation section that names the unvalidated target attribute. None of these, however, estimates the target field's error rate at scale. Since the field drives most assignments, a systematic bias there could move the identity share outside the reported boundary-sweep envelope, because that envelope conditions on the LLM target labels. The recommendation is therefore to retain the reader's CONDITIONAL verdict rather than accept as-is or reject: the central construct-validity argument is plausible and partially supported, but the decisive measurement input needs a human audit. This does not change the reader's verdict, so verdict_should_be is UNCHANGED.","tokens_in":16965,"tokens_out":6869,"duration_ms":69967,"concrete_test":"Sample ~300 gate-positive items from the 5,457, stratified by operation and script (oversample the 850 machine-translated items and the ambiguous ethnic/immigration target categories). Have two independent annotators, blind to the LLM output, code only the target-group attribute using the frozen taxonomy. Measure per-category Cohen's kappa/accuracy for the LLM target field, then recompute the 50.1/30.4/19.5 split and the 5,457 vs 2,733 overstatement using the human-corrected labels (or a misclassification-rate sensitivity bound). If per-category accuracy is high (e.g., kappa >= 0.8) and the recomputed identity share remains within 45-67% with the overstatement within 1.5-2.2x, the unvalidated-field concern fails; if identity share moves by more than about 10 percentage points, the headline percentages and the 2x claim need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 shows the auditable rule maps the frozen target-plus-narrative taxonomy to identity/partisan/geopolitical, and 69.0% of assignments are target-decisive; the narrative-diagnostic assignments (30.9%) and the 10.8% text-level override do not shield the headline split from large-scale errors in that one field. The rule-level validation is real but small: kappa=0.52 on 110 items and kappa=0.44 between two human coders on an 80-item boundary subset. It validates the whole rule against expert judgments on an enriched sample, not the target attribute's per-category accuracy across 5,457 positives—850 of them machine-translated non-English items. Section 7 concedes exactly this: 'the decisive target-group attribute is itself LLM-derived and is not separately validated at scale.' The headline 2x margin is 5,457 versus 2,733 identity-typed positives; if the characterization model systematically labels partisan attacks as ascriptive-identity attacks, or state-targeted invective as group attacks, both the 50.1/30.4/19.5 split and the 2x margin move. The Section 5.1 boundary sweep varies the identity decision rule while holding target labels fixed, so it does not bound this source of error. This is the load-bearing soft spot; it is an acknowledged empirical gap, not a logical inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 25.08M tweets from seven state-attributed influence operations and argues that widely reported 'hate' or 'toxicity' rates for such content rest on a measurement error: the detectors validate a broad construct of hostile/divisive out-group targeting rather than hate speech proper. The authors first validate a two-prompt LLM gate (κ=0.82 on a 100-item gold), then apply an auditable rule to the 5,457 gate-positive items, typing them as identity-based hate (50.1%), partisan divisiveness (30.4%), or geopolitical invective (19.5%); only 18.7% are both identity-directed and dehumanizing/inciting. They report that the broad flag therefore overstates hate by roughly 2–5×, that six of seven operations fall into three construct regimes (identity hate, geopolitical invective, partisan divisiveness), and that the divisive/hate boundary is itself unstable across expert annotators (κ=0.37–0.50) and models (best κ=0.601 against the expert majority). The paper frames the contribution as typing the construct before counting it.","tokens_in":17264,"tokens_out":4896,"duration_ms":50993,"significance":"If the quantitative composition is robust, this is a valuable measurement critique for computational social science and platform governance: it provides a transparent, auditable rule, separates the broad hostility construct from narrower hate, and shows that a single scalar 'hate rate' flattens heterogeneous operations. Strengths include the separate validation of gate and typing rule, the boundary-enriched human gold with three experts and nineteen models, the explicit boundary sweep (identity share 45.1–66.9%), the conservative two-prompt consensus, the concurrent validation on the IRA role-labeled corpus, and the clearly scoped, falsifiable transfer claim. The main reservation is that the decisive target-group attribute is LLM-derived and not validated at scale, so the exact split and overstatement margins are provisional despite the paper's careful internal validation.","major_comments":[{"comment":"The central numbers—50.1% identity, 30.4% partisan, 19.5% geopolitical, 18.7% core, and the 2× overstatement—all rest on the LLM-derived target-group attribute. The paper concedes in §7 that this attribute 'is itself LLM-derived and is not separately validated at scale,' and §4.3 reports that 69.0% of assignments are target-decisive. The §5.1 boundary sweep varies the decision rule while holding target labels fixed, so it does not bound characterization-model bias on that field. A systematic error in labeling, for example calling partisan attacks identity-based or state-targeted invective group-based, would move the 50.1/30.4/19.5 split and the 2× margin. Please validate the target field on a stratified sample of the 5,457 positives (including the 850 machine-translated items) against human gold, or provide a sensitivity analysis that perturbs target labels and recomputes the composition","section":"§5.1; Table 2"},{"comment":"The headline composition percentages are point estimates with no uncertainty despite moderate typing reliability (κ=0.52) and widely varying positive-set sizes across operations (e.g., BD-op has tens of items). Report bootstrap confidence intervals for the overall and per-operation splits, or at least clearly flag rows whose composition is statistically unstable. The 45.1–66.9% envelope in §5.1 is a sensitivity range for the rule's boundary choices, not a confidence interval for the underlying target attribute, and should not be read as covering characterization-model error.","section":"§5.1; Table 2"},{"comment":"The abstract and introduction assert that existing reported 'hate'/'toxicity' rates 'rest on a measurement error,' but the 2–5× overstatement is computed for the authors' own two-prompt Qwen gate, not for the detectors used in the cited prior work [46,45]. The transfer to prior detectors is assumed rather than directly tested. The internal comparison (5,457 vs. 2,733; 5,457 vs. 1,023) is sound for this gate, but the manuscript should either test the actual prior detectors on a shared sample or soften the framing to a conditional claim about how broad-hostility detectors behave.","section":"Abstract, §1, §6"}],"minor_comments":[{"comment":"Typo: 'shared productmanufactured divisiveness' needs a space; Table 1 caption has 'T able'. Also 'dehumanizing hardcore' is inconsistently hyphenated as 'hard-core' elsewhere.","section":"Abstract; Table 1"},{"comment":"Table 3 is dense but effective; consider adding a note that greedy decoding was used for all models and that API model sampling variability was not assessed, since the 'no model exceeds κ=0.601' claim is single-run.","section":"§5.3"},{"comment":"The concurrent validation on the Clemson IRA corpus is a useful check, but its candidate selection uses the same cue-and-target filter, so the role-level percentages are not corpus base rates; the text acknowledges this, but a one-sentence reminder in the figure caption would help readers.","section":"§5.4"},{"comment":"Reference [9] is a companion preprint by the same author; consider clarifying the relationship in one sentence to avoid any appearance of redundancy with this manuscript.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central contribution is defensible and well-scoped, but the quantitative headline is not yet fully load-bearing because the decisive target-group attribute is unvalidated and the transfer to prior detectors is assumed. The paper's own §7 candidly acknowledges the first gap, which I regard as a fixable empirical weakness rather than a fatal flaw. I would also note for the editor that the manuscript cites seven Ferrara-authored works, including two companion papers; this is not improper, but it increases the importance of ensuring the companion-paper relationship is transparent to readers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the punchline: the paper makes a real measurement point. The gate used in influence-op research, validated only as broad hostility/divisiveness, is being reported as a hate rate. On seven archive operations, the paper shows that only about half of gate-positive content is identity-directed, and only ~19% meets a stricter dehumanizing/inciting identity standard. That's a substantial over-attribution, and the finding is robust to plausible boundary choices. This is the first decomposition of an influence-op corpus into identity/partisan/geopolitical constructs, and the per-operation regimes (Russia = identity hate, Iran = geopolitical invective, Venezuela = partisan) are a useful corrective to flat 'hate' rates.\n\nThe design is genuinely careful: a two-prompt LLM gate validated against human gold (κ=0.82), a deterministic typing rule with a frozen taxonomy, expert validation (κ=0.52) within the human-human agreement band (κ=0.44), sensitivity sweeps, and a concurrent validation on a role-labeled IRA corpus. The paper is unusually candid about its own limits—Section 7 acknowledges the decisive target-group attribute is LLM-derived and not validated at scale.\n\nThe soft spot is exactly that. The typing rule is target-decisive for 69% of assignments, so a systematic mis-assignment in the LLM's target extraction (say, partisan attacks labeled as identity attacks) would shift the 50/30/20 split and the per-operation regimes. The boundary sweep in §5.1 holds target labels fixed, so it doesn't bound this error source. That said, the main over-attribution claim is more robust than the stress-test suggests: even if the target model inflates identity by mis-labeling partisan content as identity, the identity share would have to drop below ~20% to erase the 2x overstatement, which is implausible given the rule's conservative treatment of political content. The qualitative conclusion—the broad gate is not a hate detector—survives. What suffers is the precision of the composition estimates and the per-operation regime claims.\n\nOther issues are minor: the headline percentages carry no uncertainty intervals, the extrapolation to prior detectors is a prediction rather than a test (the paper says so), and the smallest operation (BD-op, 11 accounts) is rightly flagged as unstable. The kappas around 0.5 look low but are contextualized against human-human agreement on the same boundary; that's honest.\n\nWho is this for? Computational social scientists studying influence operations, platform integrity researchers, and anyone who reports toxicity/hate rates from broad classifiers. It deserves a serious referee. If I were editor, I'd send it out with a request to validate the target attribute on a human-labeled subsample and report bootstrap intervals for the composition percentages. The central argument holds; the precision needs support.","headline":"The core measurement argument holds up—the 'hate' gate overstates hate ~2x on these operations—but the precise composition split rests on an LLM-derived target field the paper admits it never validated at scale.","tokens_in":17764,"tokens_out":3582,"would_cite":true,"duration_ms":34963,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Routinely reported 'hate' and 'toxicity' rates for state-backed influence operations over-count hate by about twofold, because the detectors behind them actually measure a broader category of hostile or divisive out-group targeting.","keywords":["hate speech detection","influence operations","measurement validity","content moderation","social media manipulation","LLM annotation","manufactured divisiveness","partisan hostility"],"falsifier":"Re-code a stratified sample (e.g., 300 posts) of the 5,457 gate-positive posts with human labels for whether the target is an identity group, a partisan actor, or a state; if the human-validated identity share falls far below the model-derived share, or if human-coded hate-speech prevalence matches the broad flag, the central overstatement claim would be refuted.","tokens_in":16804,"feed_emoji":"📊","tokens_out":7526,"duration_ms":58949,"temperature":0.7,"pith_summary":"The paper argues that the widely reported 'hate' and 'toxicity' rates for state-backed influence campaigns are inflated by a measurement error: the detectors used to produce them are validated to catch a broader phenomenon—hostile or divisive attacks on an out-group—not hate speech proper. On 25 million posts from seven government-attributed campaigns, it separates the flagged content into three types and shows that under its stated criteria only about 19% of the flagged posts are both identity-directed and dehumanizing or inciting. Reporting every flagged post as 'hate' thus overstates hate roughly twofold. The mix also differs systematically by operation, with six of the seven campaigns sorting into three construct regimes, and the paper introduces 'manufactured divisiveness' as the shared product. A sympathetic reader would care because the correction changes how we measure platform harm and how we compare state actors.","feed_headline":"Hate rates overstate influence-op hostility by ~2x","feed_subtitle":"Across 25M posts from 7 state campaigns, most flagged 'hate' is partisan or geopolitical invective.","key_machinery":"The argument is carried by a two-stage instrument. Stage one is a broad-construct gate: a language model scores each (tweet, target) pair for hostile or divisive out-group targeting, and an item is positive only if two prompts (one permissive, one strict) both flag it; the gate is validated against human gold at Cohen's kappa = 0.82. Stage two is an auditable rule over a frozen 11-dimension characterization taxonomy produced by the model: the rule deterministically maps the target-group and narrative fields to one of three constructs (identity hate, partisan divisiveness, state/geopolitical invective) and flags a dehumanizing/inciting hard core. The division of labor is key: the broad gate i","core_discovery":"The central discovery is that the construct a detector measures matters: a gate validated at high agreement as a detector of hostile or divisive out-group targeting is not a hate-speech detector. Applied to 5,457 gate-positive posts across seven operations, an auditable typing rule assigns 50.1% to identity-based attacks on people, 30.4% to partisan attacks, and 19.5% to invective against states and foreign policy; only 18.7% meets the narrower standard of identity-directed dehumanizing or inciting content. Consequently a single 'hate rate' compresses categorically different content and overstates hate by a factor of about two (up to five against the strictest core). The composition is not u","pith_inferences":["The same over-attribution mechanism should apply to any study reporting a toxicity or hate rate from a detector validated against a broad hostility construct; the correction factor is set by how much of the hostile tail in that corpus is non-identity invective and can be estimated by re-typing a sample.","If the target-group attribute carries bias (the paper's unvalidated field), the 50/30/20 split is a lower bound on identity hate; the overstatement conclusion still holds within an envelope of 1.5–2.2x, so the core finding is unlikely to reverse.","Platforms and researchers could adopt the two-stage design and report composition (identity vs. partisan vs. geopolitical) alongside any rate, which would change how influence-operation harm is compared across actors.","Because even experts disagree on the boundary, automated enforcement keyed to the identity boundary should be treated as triage, not as a fixed threshold; the paper frames this as a reporting recommendation, which a reader can extend to policy."],"forward_implications":["Reported hate rates for these seven operations overstate hate speech by about 2x; measured against the narrowest defensible core the gap widens to about 5x.","The three construct regimes show that operations scoring similarly under a broad gate produce categorically different hostile content, so cross-actor comparisons based on a single rate are misleading.","The divisive/hate boundary has no annotator-stable reading: three experts agree only moderately and no model exceeds kappa 0.601 against the expert majority, so any single gold label inherits an idiosyncratic reading.","The leading Russia-attributed operation remains the only one with a non-trivial prevalence floor at the dehumanizing-identity-hate core, so its elevation persists—and becomes more distinct—under the narrowed construct.","Account-level concentration is construct-specific: one operation's high concentration under the broad gate disappears when attention is restricted to identity hate, implying that broad-construct concentration should not be read as identity-hate concentration."],"fun_headline_variants":["Influence-op 'hate' is mostly partisan or state invective","Influence-op 'hate' overcounted ~2x—only 18.7% is real hate","State ops' 'hate' is 50% identity, 30% partisan, 20% state invective","Hate detectors mislabel political invective as hate in state ops","Only 18.7% of flagged 'hate' in state ops meets the bar"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The decomposition rests on the model-derived field that identifies whom each post targets; the paper does not validate that field at scale against human labels, and 69% of its typing decisions depend on it, so any systematic error in target assignment would shift the split and the overstatement factor.","fun_headline_variants_meta":{"raw":{"variants":["Influence-op 'hate' is mostly partisan or state invective","Influence-op 'hate' overcounted ~2x—only 18.7% is real hate","State ops' 'hate' is 50% identity, 30% partisan, 20% state invective","Hate detectors mislabel political invective as hate in state ops","Only 18.7% of flagged 'hate' in state ops meets the bar"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001321,"raw_usage":{"total_tokens":5297,"prompt_tokens":906,"completion_tokens":4391,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":4283}},"tokens_in":650,"tokens_out":4391,"duration_ms":29992,"temperature":1.0,"reasoning_tokens":4283,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T01:55:09.555028+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-code a stratified sample (e.g., 300 posts) of the 5,457 gate-positive posts with human labels for whether the target is an identity group, a partisan actor, or a state; if the human-validated identity share falls far below the model-derived share, or if human-coded hate-speech prevalence matches the broad flag, the central overstatement claim would be refuted.","supporting_citations":[],"review_version":1}