{"id":"053c6d56-0345-462c-9e99-82ad5e656fb2","arxiv_id":"2505.03773","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Female computer science academics receive a higher share of threatening and severely toxic replies on Twitter, and they express stronger emotions in their posts than male academics.","lead":"This study analyzed tweets, retweets, and replies from male and female computer science academics at top 20 US universities, finding that women receive more threatening and severely toxic replies than men. The paper documents gender differences in topics, emotions, and audience responses, contributing to evidence on online scholarly communication.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reply toxicity gender gap in Table 6 is unadjusted for topic/engagement and lacks uncertainty; the central claim is not yet established.","rationale":"The reader's weakest assumption is that reply differences are attributable to author gender rather than topic, sentiment, or popularity confounds; my reading of §4.1–4.2 agrees. This is the single most load-bearing issue because the headline result is the reply toxicity gap, and Table 6 cannot sustain it without a covariate-adjusted model. The classifier experiment is sometimes presented as 'regression analysis' but is a classification task; even a well-performing classifier only shows that reply text carries information about the author's gender, which is exactly what topic leakage would produce. I am not arguing the finding is false; external literature and the raw percentages suggest a directionally plausible effect. However, for this paper, the evidence is not yet convincing. A mixed-effects model with topic and engagement covariates directly tests the concern. Since this matches the reader's identified weakness and the conditional verdict, I recommend no change to the conditional verdict.","tokens_in":14146,"tokens_out":3807,"duration_ms":40453,"concrete_test":"Fit a mixed-effects logistic regression to all replies, with outcomes for Perspective threat and severe-toxicity scores above 0.4, a fixed effect for author gender, covariates for the original tweet's topic cluster (from §4.1.1), tweet sentiment score, log(retweets+1), log(favorites+1), and reply length, and random intercepts for tweet and author. If the gender coefficient is not statistically significant or changes sign when these covariates are included, the headline claim is not supported. Report the coefficient, standard error, and cluster-robust p-value.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'female academics are more frequently subjected to severe toxic and threatening replies' is supported only by Table 6 and the §4.2 classifier experiment. Table 6 reports unadjusted percentages: replies with high threat are 15.6% for female-authored tweets versus 10.7% for male-authored tweets, and high severe toxicity is 20.4% versus 18.3%. No confidence intervals, significance tests, or effect sizes are reported, and no account is taken of the nested structure of the data (many replies per tweet, many tweets per author). The larger issue is confounding by tweet content and engagement. Section 4.1 shows the genders post different topic mixes: men tweet more about Current US Society and Opinions, Machine Learning, and Personal Thoughts; women tweet more about Engaging AI Events and Workshops. Section 4.1 also shows engagement differs by topic and author gender. Replies to a politically charged post and replies to a workshop announcement differ in toxicity for reasons unrelated to the author's gender. The BERTweet 'regression' in §4.2 does not address this: it is a binary classifier predicting the author's gender from reply text, so it can exploit topic or engagement cues (e.g., reply content about workshops versus politics) rather than gender-directed hostility. The paper's restriction to the same universities and positions equates the authors, not the tweets. As written, the evidence does not isolate gender as the driver of differential reply toxicity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates gender differences in the online scholarly discourse of computer science academics on X/Twitter, using a dataset of 627 academics from top 20 US universities. It analyzes tweets and retweets for topic prevalence, sentiment, emotion, and writing style (via an LLM), and replies for toxicity and threats using Google's Perspective API. The central claim is that replies to female academics more frequently contain severe toxic and threatening language than replies to male academics, alongside secondary claims about women's stronger emotional expression and differences in writing style. The paper is observational and descriptive, presenting comparisons of percentages and average scores without statistical inference or confounding control.","tokens_in":14396,"tokens_out":3348,"duration_ms":36826,"significance":"If the central claim were supported by rigorous evidence, this would be a valuable contribution to the literature on gender-based harassment in academic social media, with direct implications for platform moderation and academic inclusion policies. The authors have assembled a purpose-built dataset and used multiple NLP tools (topic clustering, sentiment, emotion, Perspective API), and the topic analysis includes a manual verification step, which are strengths. The study also benefits from focusing on a relatively homogeneous population (CS faculty at top universities), which mitigates some demographic confounding. However, the absence of statistical tests, the lack of control for tweet content and engagement, and the mislabeled classifier experiment currently prevent the paper from establishing its headline finding.","major_comments":[{"comment":"The central claim that female academics receive more threats and severe toxicity is based on unadjusted percentages without any measure of uncertainty. For example, the high-threat percentages are 15.6% for female-authored tweets versus 10.7% for male-authored tweets, and the severe-toxicity percentages are 20.4% versus 18.3%. No confidence intervals, significance tests, or effect sizes are reported anywhere in the paper. Moreover, the data have a nested structure — multiple replies per tweet and multiple tweets per academic — which violates the independence assumption of simple comparisons. The authors should use cluster-robust inference or a mixed-effects model that accounts for tweet and author random effects; otherwise the differences could easily be within sampling variability.","section":"Section 4.2, Table 6"},{"comment":"The analysis does not control for the content or popularity of the original tweets, which is a load-bearing confound for the main finding. Section 4.1 itself shows that male and female academics post different topic mixes (Figure 2b) and that engagement differs by topic and gender (Figure 2c). Replies to politically charged posts about 'Current US Society and Opinions' are plausibly more hostile than replies to workshop announcements, independent of the author's gender. The BERTweet experiment in Section 4.2 is labeled a 'regression analysis' but is actually a binary classifier that predicts the author's gender from reply text; it does not adjust for topic, engagement, follower count, or tweet length. As such, the classifier can exploit topic cues rather than gender-directed hostility, and the claim that the observed differences are attributable to the author's gender is not established.","section":"Sections 4.1 and 4.2; Figure 2"},{"comment":"The interpretation of the confusion matrices as evidence of gender-directed hostility is circular. Training a classifier to distinguish replies to male-authored tweets from replies to female-authored tweets and then reporting that 'the model reliably identifies threatening and toxic replies targeting women' conflates classifiability with evidence about the cause of the hostility. The classifier's accuracy could reflect any systematic difference in replies, including topic, sentiment, or engagement. To support the gender-attribution claim, the authors need to either compare replies to gendered tweets matched on topic and engagement, or explicitly test whether reply toxicity varies with author gender after controlling for tweet-level covariates. As written, the experiment does not provide the stated control.","section":"Section 4.2, Figure 6"}],"minor_comments":[{"comment":"The abstract states that women 'post slightly more' on one topic, but Figure 2b shows average counts with no indication of variability or significance; please clarify whether this difference is statistically meaningful or descriptive only.","section":"Abstract and Section 4.1.1, Figure 2b"},{"comment":"The table layout is confusing: the column labels 'Male Female' appear in both blocks, and the meaning of the '±' values in the first block is not defined (presumably standard deviation). Please reformat and define all symbols.","section":"Table 6"},{"comment":"The Mixtral-based writing-style analysis is used without validation: there are no agreement statistics with human annotations, no description of how 'Not Sure' responses were handled, and no clarification of whether the percentages are per tweet or per author. These details are needed to assess the reliability of the empathy and personal-experience claims.","section":"Section 4.1.4, Table 5"},{"comment":"The threshold for 'high perspective score' (> 0.4) is arbitrary. The authors should justify this cutoff or show that the main results are robust across a range of thresholds.","section":"Section 4.2"},{"comment":"These figures lack axis labels and sample sizes; it is unclear how many replies underlie each density curve or confusion matrix. Please add the required annotations.","section":"Figures 5 and 6"},{"comment":"The word 'significant' is used in several places without a statistical test, such as 'significantly more engagement' in Section 5 and 'gender plays a significant role' in the conclusion. Please either provide the corresponding tests or use non-statistical wording.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important and timely topic, and the dataset is a useful resource, but the current analysis does not yet support the causal-sounding language in the abstract and conclusion. The main issues — missing statistical inference and uncontrolled confounding by topic/engagement — are fixable within the scope of a revision, so I recommend major revision rather than rejection. I would also gently suggest that the authors reduce the number of self-citations (e.g., Refs. 33, 46, 58, 59) and ensure all references are complete, as this is a quick way to strengthen the paper's stance with readers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things up front. First, this is a genuinely descriptive paper: it builds a clean, hand-curated dataset of CS academics at top US universities (627 people, ~340k tweets) and documents a series of gender differences in topics, sentiment, emotion, and writing style. Second, the central claim that female academics receive more threatening and severely toxic replies is plausible but not demonstrated by the evidence as presented. The paper needs major revision, not rejection.\n\nWhat is actually new here is the population: prior work on gender and online harassment has mostly looked at politicians, journalists, or the general public, not a relatively homogeneous group of CS professors. The dataset construction is careful (manual handle matching, position and gender coding from homepages), and the topic clustering with BERTweet and manual label verification is reasonable. The descriptive findings—women express stronger positive and negative sentiment around events, women show more empathy and share personal experiences, men tweet more about AI/ML and US politics—are consistent with prior literature, which gives me some confidence the pipeline is working.\n\nThe soft spots, in proportion. The biggest one is Table 6: it reports raw percentages of replies above a toxicity threshold (threat 15.6% for women vs 10.7% for men) without confidence intervals, significance tests, effect sizes, or any adjustment for the nested structure of the data (many replies per tweet, many tweets per author). Given the gender differences in topics shown in Section 4.1—men tweet more about politics and AI, women more about workshops—the reply differences could easily be confounded by topic or engagement. The paper asserts comparability by restricting to the same universities and positions, but that equates the authors, not the tweets.\n\nThe \"regression analysis\" in Section 4.2 is mislabeled: it trains a binary classifier to predict the author's gender from the reply text. That is not a regression and does not control for anything; the classifier can learn topic cues, such as replies about workshops versus US politics, rather than gender-directed hostility. The writing-style analysis relies on Mixtral labels with no human validation, and the paper releases neither code nor data, limiting reproducibility. These are all addressable.\n\nOn the positive side, the paper is honest about its limitations (small sample, binary gender, API constraints) and its citations to prior work are extensive. The self-citations are not excessive and several are genuinely relevant.\n\nMy take: this is a serious observational study with a clear, falsifiable research question and a plausible but unproven headline finding. It deserves a serious referee. I would send it to peer review with a request for major revision: add statistical inference (including confidence intervals and appropriate multilevel models), include topic and engagement as covariates, validate the LLM labels on a human-annotated sample, and release the data and code. With those changes, the paper could make a modest but real contribution to the literature on gender and online academic discourse.","headline":"A useful descriptive study of gender patterns in CS academics' Twitter activity, but its headline claim about toxic replies is not yet established because the analysis lacks statistical inference and does not control for topic or engagement.","tokens_in":14927,"tokens_out":1689,"would_cite":false,"duration_ms":19988,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that replies to female academics contain more threats and severe toxicity than replies to male academics, even in a matched sample of computer science professors from top US universities.","keywords":["gender bias","scholarly discourse","toxicity","Twitter","harassment","sentiment analysis","academic social media","computer science"],"falsifier":"A matched-topic study of replies to male and female academics' tweets about the same news event, paper, or identical text would settle the claim: if the threat and severe-toxicity gap disappears once the tweeted content is held fixed, the gender attribution is falsified.","tokens_in":13900,"feed_emoji":"💬","tokens_out":7896,"duration_ms":73631,"temperature":0.7,"pith_summary":"This paper studies how gender shapes online scholarly discourse by analyzing Twitter/X posts and replies from computer science professors at top US universities. It finds that while male and female academics discuss largely similar topics, audiences respond differently: replies to female academics more frequently contain threats and severe toxicity, whereas identity attacks reach similar rates for both genders with a slight skew toward men. The paper also reports that women express stronger positive and negative sentiments and more empathy, while men post more about AI, machine learning, and personal perspectives. The authors intend these findings to show that gender influences both self-presentation and audience hostility in academic social media, and to motivate more inclusive and safer scholarly engagement online.","feed_headline":"Female academics on Twitter draw more threats and severe toxicity","feed_subtitle":"An analysis of 109,000 replies to CS professors finds gendered hostility.","key_machinery":"The central object is a comparative reply-toxicity analysis using the Perspective API's scores for threat, severe toxicity, and identity attack, combined with a fine-tuned BERTweet classifier that predicts the original author's gender from reply text. The Perspective scores provide a continuous measure of hostility, and the classifier tests whether reply language is sufficiently gender-distinct that replies can be assigned to the target's gender. Supporting analyses use topic clustering of tweet embeddings, sentiment and emotion classifiers, and an LLM-based writing-style questionnaire.","core_discovery":"The paper's central discovery is a gendered asymmetry in the hostility of replies directed at academics on X/Twitter. Measuring reply text with toxicity scores, the authors find that a higher percentage of replies to female academics cross a high threshold for threat (15.6% vs 10.7%) and severe toxicity (20.4% vs 18.3%) compared with replies to male academics, while the proportion of identity attacks is nearly equal (21.1% vs 22.7%). A classifier fine-tuned on reply text can reliably identify threatening and severely toxic replies aimed at women and identity attacks aimed at men, suggesting the language directed at each gender is measurably different. The paper also finds that male-authored tweets draw more engagement, female academics post with stronger positive and negative sentiment around events, and female writing style is more empathetic and personal.","pith_inferences":["A natural test of the toxicity claim would be a matched-topic design: compare replies to male and female academics' tweets about the same paper, event, or standardized prompt; if the threat and severe toxicity gap persists once content is held fixed, the gender attribution is much stronger.","The BERTweet classifier's ability to infer the target's gender from reply text could be repurposed as a low-cost auditing tool to estimate gendered harassment exposure across other fields or platforms.","Because the sample is limited to computer science professors at top-20 US universities, the findings may understate harassment in less visible or less protected academic contexts.","If platforms incorporate reply-toxicity scores into moderation, the gendered distribution found here suggests that automated systems should be tuned separately for threat and identity-attack categories rather than treated as a single toxicity bucket."],"forward_implications":["If the central claim holds, gender alone—not just content—shapes the hostility of audience responses to academics on Twitter/X.","Moderation and harassment-detection systems may need to account for the gendered distribution of threat and severe toxicity, since these reply types are more common for female academics.","The finding that identity attacks skew toward male academics at the highest intensities suggests that the form of abuse, not just its volume, differs by gender.","The writing-style differences (more empathy and personal sharing by women) and engagement gaps (male-authored tweets getting more retweets and favorites) imply that gender influences how academics present themselves and how their work circulates.","For scientific communication, the result implies that female academics face a less hospitable reply environment even within a comparatively elite and homogeneous population."],"supporting_citations":[{"why":"Provides the threat, severe toxicity, and identity-attack scores that measure reply hostility, the direct basis for the central claim.","marker":"[51]"},{"why":"Supplies the pretrained tweet language model used both to fine-tune the gender-prediction classifier and to embed texts for topic clustering.","marker":"[41]"},{"why":"Defines the top-20 university selection that bounds and makes comparable the academic sample.","marker":"[11]"}],"fun_headline_variants":["Female academics get more threats in Twitter replies","Gendered abuse: female CS professors face more toxic replies","Study finds female academics receive more threatening replies","Female professors on X face more severe toxicity in replies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central argument depends on the assumption that the higher threat and severe-toxicity rates in replies to female academics are caused by the author's gender rather than by differences in what they tweet about, how popular their posts are, or the topics that trigger hostile replies.","fun_headline_variants_meta":{"raw":{"variants":["Female academics get more threats in Twitter replies","Gendered abuse: female CS professors face more toxic replies","Study finds female academics receive more threatening replies","Female professors on X face more severe toxicity in replies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1390,"prompt_tokens":929,"completion_tokens":461,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":401}},"tokens_in":545,"tokens_out":461,"duration_ms":5423,"temperature":1.0,"reasoning_tokens":401,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:11:33.411708+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A matched-topic study of replies to male and female academics' tweets about the same news event, paper, or identical text would settle the claim: if the threat and severe-toxicity gap disappears once the tweeted content is held fixed, the gender attribution is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the threat, severe toxicity, and identity-attack scores that measure reply hostility, the direct basis for the central claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the pretrained tweet language model used both to fine-tune the gender-prediction classifier and to embed texts for topic clustering."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the top-20 university selection that bounds and makes comparable the academic sample."}],"review_version":1}