{"id":"b08ad083-726a-4823-96a2-cab886e7dd0f","arxiv_id":"2505.07212","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Russian, Iranian, and Chinese influence operations on Twitter differ systematically in the sentiment, emotion, and toxicity of their English-language tweets.","lead":"This paper analyzes two million English tweets from Twitter's public datasets of state-sponsored influence operations linked to China, Iran, and Russia, using sentiment, emotion, and toxicity classifiers. It reports that Russian campaigns are the most negative and toxic, Chinese campaigns are the most positive and neutral, and Iranian campaigns sit in between, suggesting tailored messaging strategies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Iran-18 subset is machine-translated and all classifiers are English/Western-trained; without validation, the cross-country differences in Tables 2-4 may be partly measurement artifacts.","rationale":"The central claim is a comparative statement about measured sentiment, emotion, and toxicity. All measurements flow through two off-the-shelf English classifiers. The one dataset with non-English origin—Iran-18—is exactly where the pipeline diverges through translation, and it is also where the abstract's 'blend' characterization is most dependent. The authors themselves flag Western-discourse training and the absence of error analysis as limitations, so this is not an external-consensus disagreement but an internal validity threat to the measurement. The reader identified the same weakest assumption, and I agree. The concern is substantial but testable; the conditional verdict already requires additional validation, so this review does not move the verdict.","tokens_in":9944,"tokens_out":4974,"duration_ms":54806,"concrete_test":"Take a stratified random sample of 1,000 tweets from each of the four datasets (for Iran-18, retain the original Persian text and the English translation used in the pipeline). Have at least two trained annotators label sentiment (negative/neutral/positive), emotion, and toxicity for each tweet, working on original text where available or on a verified translation. Compare TweetNLP and Perspective predictions against these gold labels per dataset; then recompute Tables 2-4 using only tweets whose automatic label agrees with gold labeling, or apply a debiasing correction. If the Russia > Iran > China ordering in negative sentiment and toxicity survives, the measurement concern is mitigated; if classifier accuracy is markedly worse for Iran-18 than for China-19 or Russia-19, the central comparisons are not currently interpretable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 states that, because the 2018 Iran dataset had few English tweets, the authors translated tweets into English; Section 4's Algorithm 1 has no translation step, and the translation method/model is unspecified. One quarter of the 2M analyzed tweets therefore enters the pipeline in a fundamentally different form. TweetNLP and Perspective API are English-only models trained primarily on Western discourse (the authors concede in Section 6 that this limits cross-cultural accuracy). Machine-translated Persian ('translationese') can systematically shift sentiment/emotion/toxicity scores, so the Iran-18 row in Tables 2-4—and the abstract's characterization of Iran as blending antagonistic and supportive tones—may reflect translation artifacts rather than Iranian strategy. No error analysis or per-language validation is reported. Russia's 9x toxicity over China and 3.5x over Iran is large, but the comparative claim treats all four rows as commensurable; if Iran-18 labels are biased, the 'distinct patterns' inference is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes sentiment, emotion, hate speech, offensive language, and toxicity in 2 million tweets from state-sponsored influence operations (SIOs) by China, Iran, and Russia, using Twitter's publicly released datasets. The authors apply TweetNLP and Perspective API, report per-country proportions in Tables 2–4, and claim that Russian campaigns are predominantly negative and toxic, Iranian operations blend antagonistic and supportive tones, and Chinese activities emphasize positive and neutral rhetoric. The paper is a short descriptive study with no statistical inference or validation of the external classifiers on the target languages.","tokens_in":10128,"tokens_out":4781,"duration_ms":45197,"significance":"If the central claim of distinct affective and rhetorical patterns per country were robust, this would be a useful descriptive contribution to the literature on state-sponsored influence operations, leveraging publicly available Twitter datasets and reproducible off-the-shelf tools. The paper's strengths include a large sample (2M tweets), transparent counts in the tables, and the use of established classifiers (TweetNLP, Perspective API). However, the headline cross-country comparisons are based on raw proportions without confidence intervals or significance tests, one of the four datasets (Iran-18) enters the pipeline as machine-translated text with the translation procedure unspecified, and the classifiers are acknowledged to be trained on Western discourse. These issues directly affect the validity of the 'distinct patterns' claim as currently stated.","major_comments":[{"comment":"The Iran-18 tweets were machine-translated into English, but the translation method/model is never specified, and Algorithm 1's pipeline contains no translation step. Because TweetNLP and Perspective API are English-only models, the Iran-18 row in Tables 2–4 may reflect translation artifacts (e.g., 'translationese' systematically shifting sentiment, emotion, or toxicity scores) rather than the actual Iranian strategy. Since Iran-18 constitutes one quarter of the analyzed data and feeds directly into the abstract's characterization of Iran as blending antagonistic and supportive tones, the authors must either specify the translation procedure, validate the classifiers on the translated text, or provide a robustness analysis that excludes or re-weights Iran-18.","section":"Section 3 and Algorithm 1"},{"comment":"The headline comparisons (e.g., 'Russian IOs contain more than 9 times toxic content than Chinese IOs' in the Introduction, and the cross-country proportionality claims throughout Section 5) are made from samples of 500,000 tweets per dataset with no confidence intervals, standard errors, or significance tests. While the large sample sizes imply tiny sampling variability under simple random sampling, the samples are not independent across time, and the measurement pipeline itself is a source of error. The claim of 'distinct patterns' requires at least a bootstrap confidence interval or a formal comparison (e.g., chi-square or proportion tests) to establish that the observed differences exceed what could arise from sampling or model noise. Without this, the central inference is not supported beyond descriptive observation.","section":"Section 5, Tables 2–4"},{"comment":"The four datasets cover different time periods (China Feb 2008–Aug 2019, Iran-18 Dec 2010–Aug 2018, Iran-19 Jul 2017–Mar 2018, Russia Aug 2010–Nov 2018). The paper argues that random sampling 'does not introduce temporal misalignment, as the selection was not time-bound,' but this only ensures that sampling is not time-selected; it does not address the inherent non-overlap of the operating periods. Cross-country differences in sentiment or toxicity could therefore reflect temporal shifts in platform policies, world events, or campaign objectives rather than stable country-level strategies. The authors should either restrict the comparison to a common time window, or explicitly discuss and control for temporal confounding in the interpretation of Tables 2–4.","section":"Section 3 and Table 1"},{"comment":"The external classifiers were trained primarily on Western/English discourse, and the paper itself acknowledges in Section 6 that this 'limits the model's ability to accurately capture abusive content across different cultural contexts.' No per-language validation, error analysis, or calibration against human labels is provided, despite the explicit statement that the study 'did not perform a comprehensive error analysis.' For a comparative claim across languages and cultures, this is a load-bearing gap: the measured cross-country differences in hate speech, toxicity, and emotion could be artifacts of differential model accuracy rather than genuine differences in content. At minimum, the authors should report a small manual validation set for each country or compare against existing benchmarks for the tools on non-Western text.","section":"Section 4 and Section 6"},{"comment":"The claim that 'Russian IOs accumulate almost twice the amount of hate speech and offensive tweets compared to the IOs distributed by both Iran and China' is directly contradicted by Table 3: Russia has 89,220 hate-speech tweets versus China's 8,668 (a ratio of approximately 10.3), and versus Iran-18's 36,175 (ratio 2.5) and Iran-19's 40,824 (ratio 2.2). The 'almost twice' wording is only approximately correct for Iran-19 and is incorrect for China. This internal inconsistency in the interpretation of the results should be corrected, and the discussion should present the actual ratios rather than a misstated generalization.","section":"Section 6, Discussion"}],"minor_comments":[{"comment":"The description of the sampling procedure is internally inconsistent: the text first states 'we randomly selected 500 thousand English tweets from each of the four different IOs' and then says that for Iran-18, 'due to the limited number of English tweets... we translated the available tweets into English.' Please clarify the exact procedure: were 500k tweets sampled first and then translated, or were all available tweets translated and then a sample drawn?","section":"Section 3"},{"comment":"Algorithm 1 should include the translation step for non-English input, or the pseudocode should note that translation is performed as part of preprocessing for the Iran-18 dataset.","section":"Algorithm 1"},{"comment":"In the toxicity analysis, the phrase 'more than 12% increase' and '7% rise' for Iran-19 relative to Iran-18 are ambiguous: specify whether these are percentage-point differences or relative percentage changes. The table counts imply relative increases of approximately 12.8% for hate speech and 7.8% for offensive language, so the text should be explicit.","section":"Section 5"},{"comment":"There are typos: 'Chinease' should be 'Chinese' in the opening of Section 5, and 'contnet' should be 'content' in the hate speech subsection. Additionally, 'offensive-positive' is an unusual phrasing; consider using 'offensive = true' or defining the term clearly.","section":"Section 5"},{"comment":"For the word cloud analysis, the methodology for combining hate speech, offensive language, and toxic content into 'abusive speech' is not described. Please state the exact criteria (e.g., union of tweets flagged by any of the three models) and whether the word clouds are based on the 2019 datasets only.","section":"Section 5 and Figure 1"},{"comment":"The claim that Twitter published 'over 141 information operation datasets' lacks a citation; please add a reference to the Twitter data archive or the specific release page.","section":"Section 3"},{"comment":"The rationale that the 2019 dataset is 'substantially larger' is not uniformly true: Russia-19 has 920,761 tweets, which is smaller than Iran-18's 1,122,936. Please revise the justification for choosing the 2019 datasets.","section":"Section 3, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a short conference paper with a clear pipeline and potentially interesting descriptive results. The main scientific risk is that the central claim of 'distinct patterns' is stated as a finding when the analysis lacks statistical inference and contains an unspecified machine-translation step for one of four datasets. These are fixable within the scope of a revision: adding confidence intervals or significance tests, clarifying the translation procedure, adding a robustness check for Iran-18, and correcting the overstatement in Section 6. I also note that the paper does not engage with prior content-based predictive work on these same datasets (e.g., Alizadeh et al., 2020, cited as [3]), which could strengthen the framing by explicitly distinguishing the current descriptive contribution from earlier predictive models."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a straightforward, transparent descriptive study. It takes Twitter's public IO datasets for China, Iran, and Russia, runs off-the-shelf sentiment/emotion/toxicity classifiers over 500k-tweet random samples, and reports the proportions. The three-way profile—Russia mostly negative and toxic, China mostly positive/neutral, Iran somewhere in between—is a useful descriptive benchmark that I don't think appears in exactly this form elsewhere. The paper is honest about its limitations and makes no pretense of fitting a model, so the circularity burden is close to zero. That's real value, and the result is reproducible in principle since the datasets and tools are public.\n\nThe soft spots are exactly where the reader puts them. All comparisons are raw proportions with no confidence intervals and no significance tests; with 500k tweets per group, even trivial differences would be significant, so the lack of error bars matters less for the existence of a difference than for its size, but the paper makes strong claims about 'distinct strategies' from numbers that could move with a different sample or classifier. The datasets span different years (China 2008–2019, Russia 2010–2018, Iran-19 2017–2018), so the country contrast is confounded with time. And the Iran-18 dataset was machine-translated into English, with the translation method and model unspecified and no translation step in Algorithm 1. TweetNLP and Perspective are English/Western-trained, so the Iran-18 row—and the abstract's 'blend of antagonistic and supportive tones'—could be partly a translation artifact. The paper explicitly admits it did no error analysis and that the models are Western-centric, so this isn't a hidden flaw, but it does cap how much weight the comparative claims can carry. The Russia-vs-China gap (9x toxicity) is large and probably robust; the Iran row is the fragile one.\n\nWho is this for? Researchers building or evaluating IO-detection systems who want a quick descriptive baseline on the public Twitter datasets. It doesn't move the methodological needle, but it gives a clean, citable summary. I'd send it to peer review—the topic matters, the data are public, and the limitations are fixable with a more careful statistical treatment and translation validation. With a workshop-level venue it's fine as is; for a stronger venue, ask for matched baselines and confidence intervals.","headline":"A useful descriptive three-way comparison of sentiment and toxicity in state-sponsored IO tweets, but the headline contrasts outrun the statistics and one machine-translated dataset could be skewing the Iran row.","tokens_in":10623,"tokens_out":2739,"would_cite":true,"duration_ms":25620,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Russian influence tweets are nine times more toxic than China's.","keywords":["state-sponsored influence operations","sentiment analysis","emotion analysis","hate speech","toxicity","TweetNLP","Perspective API","Twitter datasets"],"falsifier":"Manually annotate a stratified random sample of, say, 3,000 tweets from each operation for sentiment, emotion, and toxicity in the original languages, with professional translation for Persian; if the Russian-negative versus Chinese-positive versus Iranian-mixed pattern does not emerge in the human labels, the measured differences are classifier artifacts rather than properties of the campaigns. A narrower check: run the same pipeline on original Persian tweets and on their English machine translations, and if the Iranian profile changes substantially across versions, the translation step is responsible for part of the result.","tokens_in":9747,"feed_emoji":"📊","tokens_out":6753,"duration_ms":59789,"temperature":0.7,"pith_summary":"This paper argues that the state-sponsored influence operations run by Russia, Iran, and China on Twitter have measurably different emotional signatures: Russian tweets are predominantly negative and toxic, Chinese tweets are mostly neutral or positive, and Iranian tweets blend antagonistic and supportive tones. The evidence comes from two million English tweets sampled from Twitter's publicly released state-affiliated account datasets, labeled with standard NLP tools for sentiment, emotion, hate speech, and toxicity. If the pattern is real, then influence operations are not a single tactic but a menu of calibrated content strategies that differ by national objective, which matters for how platforms triage and counter them.","feed_headline":"Russian influence tweets are 9x more toxic than China's","feed_subtitle":"Analysis of 2M state-sponsored tweets shows Russia polarizes, China promotes, Iran mixes tone.","key_machinery":"Two off-the-shelf classifiers carry the analysis: TweetNLP, which assigns each tweet a sentiment (negative, neutral, positive), an emotion (anger, joy, optimism, sadness), and binary hate-speech and offensive-language flags; and Google's Perspective API, which scores six toxicity dimensions (toxic, severe toxic, profanity, identity attack, insult, threat). The paper's entire argument is the cross-country comparison of the label distributions these tools produce, tabulated in Tables 2–4; no new model or linguistic theory is introduced.","core_discovery":"On equal-size random samples of 500,000 tweets per operation, the paper reports sharp cross-country contrasts. Russian operators produced the highest negative sentiment (41.2%), the most anger (57.4%), the most hate speech (17.84%), and the highest toxicity (10.69% by Perspective API's overall score), with toxic content more than nine times that of Chinese operations and about 3.5 times that of Iranian operations. Chinese operations were dominated by neutral (59.0%) and positive (34.8%) sentiment and by joy (68.1%). Iranian operations sat in between, and their 2019 dataset shows a clear escalation from 2018: anger rose from 25.6% to 40.1% of tweets, and every Perspective toxicity category increased, suggesting a shift toward more confrontational messaging. The authors read these differences as content strategies tailored to each state's geopolitical ends.","pith_inferences":["A direct test of the classifier-risk concern would be to human-annotate a few thousand sampled tweets from each operation and compare the human labels with the machine labels; if the cross-country gaps vanish under human annotation, the paper's conclusion is an artifact of the tools.","Because the two Iranian datasets (2018 and 2019) cover different campaigns and periods, the observed escalation could reflect different operations rather than a strategic shift within a single campaign; comparing the same tactic over time within one operation would separate these explanations.","The word clouds hint that each operation attacks specific targets (Chinese attacks on an exiled businessman, Russian attacks on U.S. political figures), so a natural extension is to model target entities directly rather than only aggregate tone.","The paper's focus on English tweets leaves open whether these national styles hold in other languages; comparing each operation's English and native-language output would test whether the tone is a property of the state or of the translated, foreign-facing channel."],"forward_implications":["If the claim holds, toxicity rate is a usable triage signal: Russian operations produce roughly nine times more toxic content than Chinese ones, so accounts that spike on existing toxicity APIs deserve priority review.","The 2018-to-2019 Iranian escalation across all six Perspective toxicity categories suggests that analysts should treat an operation's content profile as time-varying, not fixed.","The three national profiles give platform researchers a content-based typology — polarizing, dual-toned, and image-polishing — that can be tested on newer datasets, such as later Twitter/X influence-operation releases.","Russian reliance on anger and identity attacks implies that counter-influence efforts should focus on de-escalating emotional arousal rather than only correcting factual claims."],"supporting_citations":[{"why":"Supplies the toxicity, severe-toxicity, profanity, identity-attack, insult, and threat labels that produce the cross-country toxicity comparisons in Table 4.","marker":"[1]"},{"why":"Supplies the sentiment, emotion, hate-speech, and offensive-language labels that produce the central contrasts in Tables 2 and 3.","marker":"[8]"},{"why":"Establishes the prior result that content-based features can distinguish state influence operations on Twitter, motivating the content-analysis approach and the use of Twitter's public SIO datasets.","marker":"[3]"}],"fun_headline_variants":["Russian state tweets 9x more toxic than China's","State influence ops: Russia hates, China smiles, Iran mixes","How Russia, China, Iran weaponize sentiment in tweets","2M state tweets reveal: Russia toxic, China upbeat, Iran volatile","Russian propaganda tweets: more hate, less joy than China's"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the off-the-shelf classifiers — trained largely on Western English discourse — produce valid, comparable labels for tweets in different dialects and for machine-translated Persian content, so that the measured cross-country differences reflect actual content strategy rather than model or translation artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Russian state tweets 9x more toxic than China's","State influence ops: Russia hates, China smiles, Iran mixes","How Russia, China, Iran weaponize sentiment in tweets","2M state tweets reveal: Russia toxic, China upbeat, Iran volatile","Russian propaganda tweets: more hate, less joy than China's"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000946,"raw_usage":{"total_tokens":4035,"prompt_tokens":936,"completion_tokens":3099,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":3013}},"tokens_in":552,"tokens_out":3099,"duration_ms":20927,"temperature":1.0,"reasoning_tokens":3013,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:21:19.825934+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Manually annotate a stratified random sample of, say, 3,000 tweets from each operation for sentiment, emotion, and toxicity in the original languages, with professional translation for Persian; if the Russian-negative versus Chinese-positive versus Iranian-mixed pattern does not emerge in the human labels, the measured differences are classifier artifacts rather than properties of the campaigns. A narrower check: run the same pipeline on original Persian tweets and on their English machine translations, and if the Iranian profile changes substantially across versions, the translation step is responsible for part of the result.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the sentiment, emotion, hate-speech, and offensive-language labels that produce the central contrasts in Tables 2 and 3."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prior result that content-based features can distinguish state influence operations on Twitter, motivating the content-analysis approach and the use of Twitter's public SIO datasets."}],"review_version":1}