REVIEW 3 major objections 5 minor 1 cited by
Analyzing Political Discourse on Discord during the 2024 U.S. Presidential Election
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Republican-aligned Discord servers were significantly more toxic than Democratic-aligned ones, and sexist hate speech rose after Kamala Harris entered the race.
desk verdict First large-scale look at political Discord, with a real dataset, but the headline toxicity comparison leans on an unvalidated Twitter-trained classifier and tiny server counts. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The analysis uses three tools: political valence, a normalized difference in term usage between Republican and Democratic servers; the Word Embedding Association Test (WEAT), which measures whether target words sit closer to positive or negative reference words in Word2Vec embedding spaces trained separately on each alignment; and a RoBERTa-based multi-class hate speech classifier trained on Twitter data, which labels messages as sexism, racism, disability, sexual orientation, religion, other, or not hate. The toxicity comparison is carried by the classifier, while the valence and embedding analyses carry the topical and implicit-bias claims.
What would settle it
Manually label a random sample of messages from Republican-aligned servers before and after July 21, 2024, focusing on ones the classifier flags as sexist; if human raters see no increase in genuinely sexist content, the toxicity spike claim is falsified. More broadly, if a hand-coded sample of Discord messages yields a different hate-speech distribution than the classifier's labels, the toxicity conclusions would need revision.
Extended reading notes
Core claim
The central claim, stated in the conclusion, is that Republican-aligned servers were significantly more toxic than Democratic-aligned ones, with sexist hate speech increasing after Kamala Harris's nomination. The authors also establish, through political valence and embedding analysis, that the two sides talked about different things: Republican-aligned servers concentrated on economic policy and libertarian themes, while Democratic-aligned servers emphasized equality, LGBTQIA+ rights, and the Israel-Palestine conflict. The toxicity finding rests on a RoBERTa-based multi-class hate speech classifier, and the paper reports that sexism rose on Republican servers from the moment Harris entered the race, while racism did not.
Load-bearing premise
The hate speech findings assume that a RoBERTa classifier trained on Twitter text correctly identifies hate speech categories when applied to Discord messages, and the paper offers no Discord-specific validation of that transfer.
Editorial extensions
If this is right
- If the toxicity finding holds, Discord's decentralized moderation is allowing political communities with measurably higher hate speech to operate largely unchecked.
- The observed sexism spike after a Black woman became the Democratic nominee, with no corresponding racism spike, suggests gender bias may be a distinct driver of hate speech in such spaces.
- Discord, not just X or Reddit, is now a measurable platform for studying political engagement, and the released dataset enables follow-up work.
- The valence results imply that server-level ideological differences are visible even in short, informal chat messages.
Reading between the lines
- The lack of a racism spike alongside the sexism spike could mean the classifier undercounts racial hate speech in this context, or that gender was the salient trigger; a manual audit would tell which.
- The authors' claim that Twitter text 'closely resembles' Discord is testable: compare classifier agreement on a sample of Discord messages labeled by both the model and human raters, and by raters familiar with Discord slang.
- A natural extension is to check whether the same sexism dynamic appears in other language communities or in private servers, which the public-server sample cannot capture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a descriptive study of political discourse on Discord during the 2024 U.S. presidential election. The authors collect over 30 million messages from 81 public political servers, manually classify servers as Republican-aligned, Democratic-aligned, or unaligned (with Fleiss' kappa of 0.8191), and analyze the data with three methods: a political valence metric over terms and URLs, Word2Vec embeddings probed with the Word Embedding Association Test (WEAT), and a RoBERTa-based multi-class hate speech classifier trained on Twitter data. The main empirical claims are that Republican-aligned servers were 'significantly more toxic' than Democratic-aligned servers, that sexist hate speech rose notably after Kamala Harris's nomination, that Republican discourse emphasized economic topics while Democratic discourse emphasized equality-related topics, and that the dataset reveals alignment-specific semantic biases.
Significance. If the toxicity findings were robust, this would be a valuable contribution as one of the first large-scale political analyses of Discord, an understudied and rapidly growing platform. The paper's strengths include a publicly released anonymized dataset with a DOI, explicit manual labeling of server alignment with high inter-rater agreement, temporal segmentation around campaign events, and analysis of multiple hate speech categories rather than a single toxicity score. The valence and embedding results are plausible and presented descriptively. However, the central claim of a 'significant' toxicity gap and the sexism spike currently rest on an unvalidated cross-platform classifier and on pooled message counts without statistical inference, so the headline finding cannot be accepted as stated. The manuscript is likely repairable with additional validation and statistical analysis, but the current version is not ready to support the strong conclusion.
major comments (3)
- [§4.2 and §5.3] The entire hate speech analysis relies on a RoBERTa classifier trained on Twitter data (Antypas & Camacho-Collados 2023), with the only transfer justification being the assertion that Twitter text 'closely resembles the Discord environment.' No Discord-specific validation is reported: there is no per-category precision/recall, no manually labeled Discord sample, no calibration analysis, and no error inspection. This is consequential because the reported rates are extremely small—the largest category rate in Figure 5 is 0.55% for sexism—so a differential false-positive rate of even a few tenths of a percent, driven by the alignments' very different vocabularies (e.g., 'retarded' on the Republican side; 'feminism' and 'pronouns' on the Democratic side), could determine the sign and size of the reported gap. The authors should validate the classifier on a labeled sample of Discord messages or otherwise provide evidence that classification performance is comparable across both alignments and across the temporal periods.
- [§6 and §5.3] The conclusion that Republican-aligned servers were 'significantly more toxic' is not supported by any statistical test, confidence interval, or per-server variance estimate. The analysis in §5.3 pools all messages within each alignment, but the sample contains only 7 Republican and 9 Democratic servers; with such small server-level sample sizes, the pooled difference could easily be driven by a single outlier server. The authors should report server-level variability and conduct an appropriate test (e.g., a permutation test or a mixed-effects model with server as a random effect) before using the word 'significantly.' The same lack of inference applies to the claim of a sexism spike in the Kamala vs. Trump period.
- [§5.3.2 and Figure 1] The sexism spike after Kamala Harris's nomination is confounded with a topic shift: after July 21, both alignments dramatically increase their mentions of Harris (Figure 1), so the classifier may be flagging messages about a woman candidate as sexist regardless of the actual intent or content. The paper does not control for the volume of Harris-related discussion, nor does it test whether the classifier's false-positive rate varies with the presence of female candidate names. The authors should conditionalize the hate speech analysis (e.g., report rates excluding messages that mention Harris, or stratify by whether a message mentions her) to demonstrate that the sexism increase reflects a change in discourse rather than a compositional artifact.
minor comments (5)
- [§5.3, first sentence] The word 'dateset' is a typo for 'dataset'.
- [§5.1.4] The phrase 'it's Valence values' should be 'its valence values'.
- [§6, first paragraph] The phrase 'non political correct insults' should be 'non-politically-correct insults.'
- [§5.1.1] The sentence 'which is logic since he was then running' should read 'which is logical'.
- [Figure 5] The radar charts are difficult to compare across categories and periods; a grouped bar chart with explicit values would be clearer, especially given that the rates are under 1%.
Circularity Check
No circularity identified; the study is descriptive and its central quantities are computed from external classifiers and data-driven summaries rather than fitted to reproduce its conclusions.
full rationale
This paper is a descriptive observational study, not a derivation or prediction exercise, and I found no circular step that reduces its claims to its own inputs. The hate speech results are produced by a pre-trained RoBERTa classifier (Antypas and Camacho-Collados, 2023) trained on an external Twitter corpus; no parameter of that classifier is fitted to the Discord data or to the study's conclusions, so the toxicity and sexism findings are not circular, however questionable their cross-platform validity may be. Server alignment is assigned manually from server descriptions with reported inter-rater agreement, independently of the measured outcome variables. The political valence metric is a descriptive normalization of term frequencies across the two alignment groups, and the WEAT/embedding analyses summarize associations in the fitted embeddings rather than being constructed to force a particular result. The cited prior works are methodology sources or related datasets, not self-citations that carry the central claim. Concerns about classifier transfer, tiny base rates, and the absence of significance tests are validity or robustness issues, not circularity. The paper's limitations section candidly acknowledges coverage and demographic constraints, and nothing in the derivation chain equates an input with an output by construction.
Assumptions & free parameters
free parameters (3)
- neutral valence interval boundary =
0.33
- start of voting period =
September 20, 2024
- number of reference words in WEAT =
3 positive, 3 negative
assumptions (5)
- domain assumption The RoBERTa hate speech classifier trained on Twitter data transfers to Discord messages.
- domain assumption Server names and descriptions reflect the political alignment of the community.
- domain assumption The Discord API returned complete message history for the study period.
- domain assumption The pre-trained Word2Vec model is a politically neutral baseline.
- domain assumption Word embeddings trained on small segmented corpora are statistically meaningful.
Cite this review
Pith. "Pith review of Analyzing Political Discourse on Discord during the 2024 U.S. Presidential Election." pith.science (2026). https://pith.science/paper/XEMX2AK2
@misc{pith2026250203433,
author = {Pith},
title = {Pith review of: Analyzing Political Discourse on Discord during the 2024 U.S. Presidential Election},
year = {2026},
howpublished = {\url{https://pith.science/paper/XEMX2AK2}},
note = {Machine review of arXiv:2502.03433}
}
read the original abstract
Social media networks have amplified the reach of social and political movements, but most research focuses on mainstream platforms such as X, Reddit, and Facebook, overlooking Discord. As a rapidly growing, community-driven platform with optional decentralized moderation, Discord offers unique opportunities to study political discourse. This study analyzes over 30 million messages from political servers on Discord discussing the 2024 U.S. elections. Servers were classified as Republican-aligned, Democratic-aligned, or unaligned based on their descriptions. We tracked changes in political conversation during key campaign events and identified distinct political valence and implicit biases in semantic association through embedding analysis. We observed that Republican servers emphasized economic policies, while Democratic servers focused on equality-related and progressive causes. Furthermore, we detected an increase in toxic language, such as sexism, in Republican-aligned servers after Kamala Harris's nomination. These findings provide a first look at political behavior on Discord, highlighting its growing role in shaping and understanding online political engagement.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Who Leads in the Shadows? ERGM and Centrality Analysis of Congressional Democrats on Bluesky
Using follows, mentions, and reposts on Bluesky, the authors identify members such as Marcy Kaptur as structurally central despite low public profiles.
Reference graph
Works this paper leans on
-
[1]
Fatimah Alkomah and Xiaogang Ma. 2022. A Literature Review of Textual Hate Speech Detection Methods and Datasets. Information 13, 6 (2022). https: //doi.org/10.3390/info13060273
-
[2]
Hunt Allcott, Matthew Gentzkow, Winter Mason, Arjun Wilkins, Pablo Barberá, Taylor Brown, Juan Carlos Cisneros, Adriana Crespo-Tenorio, Drew Dimmery, Deen Freelon, Sandra González-Bailón, Andrew M. Guess, Young Mie Kim, David Lazer, Neil Malhotra, Devra Moehler, Sameer Nair-Desai, Houda Nait El Barj, Brendan Nyhan, Ana Carolina Paixao de Queiroz, Jennifer...
work page 2024
-
[3]
Dimosthenis Antypas and Jose Camacho-Collados. 2023. Robust Hate Speech Detection in Social Media: A Cross-Dataset Empirical Evaluation. In The 7th Workshop on Online Abuse and Harms (WOAH) . 231–242
work page 2023
-
[4]
Yan Aquino, Pedro Bento, Arthur Buzelin, Lucas Dayrell, Samira Malaquias, Caio Santana, Victoria Estanislau, Pedro Dutenhefner, Guilherme H. G. Evangelista, Luisa G. Porfírio, Caio Souza Grossi, Pedro B. Rigueira, Virgilio Almeida, Gisele L. Pappa, and Wagner Meira Jr. 2025. Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024). ar...
arXiv 2025
-
[5]
Ashwin Balasubramanian, Vito Zou, Hitesh Narayana, Christina You, Luca Luceri, and Emilio Ferrara. 2024. A Public Dataset Tracking Social Media Discourse about the 2024 U.S. Presidential Election on Twitter/X. arXiv:2411.00376 [cs.SI] https://arxiv.org/abs/2411.00376
arXiv 2024
-
[6]
Pedro Bento, Arthur Buzelin, Yan Aquino, Isis Carvalho, Pedro Dutenhefner, Lucas Dayrell, Caio Santana, Victoria Estanislau, Gisele Pappa, Debora Miranda, Virgilio Almeida, and Wagner Meira Jr. 2024. Impacto da Pandemia na Discussão 7https://dis.gd/discord-developer-policy 8https://discord.com/terms 9https://zenodo.org/records/14807501 10https://discord.c...
-
[7]
George R Boynton and Glenn W Richardson Jr. 2016. Agenda setting in the twenty-first century. New Media & Society 18, 9 (2016), 1916–1934
work page 2016
-
[8]
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science 356, 6334 (2017), 183–186
2017
Show all 40 references
-
[9]
Michael Conover, Jacob Ratkiewicz, Matthew Francisco, Bruno Gonçalves, Filippo Menczer, and Alessandro Flammini. 2011. Political polarization on twitter. In Proceedings of the international aaai conference on web and social media , Vol. 5. 89–96
2011
-
[10]
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. In Proceedings of the international AAAI conference on web and social media , Vol. 11. 512–515
2017
-
[11]
Emilio Ferrara, Herbert Chang, Emily Chen, Goran Muric, and Jaimin Patel. 2020. Characterizing social media manipulation in the 2020 US presidential election. First Monday (2020)
2020
-
[12]
Joseph L. Fleiss. 1971. Measuring Nominal Scale Agreement Among Many Raters. Psychological Bulletin 76, 5 (1971), 378–382. https://doi.org/10.1037/h0031619
1971 doi
-
[13]
Felipe Freire, Henrique Coelho, g1 Rio, and TV Globo. 2024. Justiça condena homem que criou grupo no Discord para estupro de vulnerável. g1 (2024). https://g1.globo.com/rj/rio-de-janeiro/noticia/2024/07/04/justica- condena-homem-grupo-discord-estupro-vulneravel.ghtml
2024
-
[14]
Thomas Fujiwara, Karsten Müller, and Carlo Schwarz. 2024. The effect of social media on elections: Evidence from the united states. Journal of the European Economic Association 22, 3 (2024), 1495–1539
2024
-
[15]
Aoife Gallagher, Ciaran O’Connor, Pierre Vaux, Elise Thomas, and Jacob Davey
-
[16]
Jacob Groshek and Karolina Koc-Michalska. 2017. Helping populism win? Social media use, filter bubbles, and support for populist presidential candidates in the 2016 US election campaign. Information, Communication & Society 20, 9 (2017), 1389–1407
2017
-
[17]
Fake news
Richard Gunther, Paul A Beck, and Erik C Nisbet. 2019. “Fake news” and the defection of 2012 Obama voters in the 2016 presidential election. Electoral studies 61 (2019), 102030
2019
-
[18]
Alicja Hagopian. 2024. JD Vance is now the least popular VP candi- date in modern history – even below Sarah Palin. The Independent (2024). https://www.independent.co.uk/news/world/americas/us-politics/jd- vance-poll-vp-popularity-sarah-palin-b2598125.html
2024
-
[19]
Daniel G Heslep and PS Berge. 2024. Mapping Discord’s darkside: Distributed hate networks on Disboard. new media & society 26, 1 (2024), 534–555
2024
-
[20]
Sam Graham-Felsen Hughes, Kate Allbright-Hannah, Scott Goodstein, Steve Grove, Randi Zuckerberg, Chloe Sladden, and Brittany Bohnet. 2010. Obama and the power of social media and technology. The European Business Review 16 (2010), 21
2010
-
[21]
Emily K Johnson and Anastasia Salter. 2022. Embracing discord? The rhetorical consequences of gaming platforms as classrooms. Computers and Composition 65 (2022), 102729
2022
-
[22]
Marcelo Sartori Locatelli, Josemar Caetano, Wagner Meira Jr., and Virgilio Almeida. 2022. Characterizing Vaccination Movements on YouTube in the United States and Brazil. In Proceedings of the 33rd ACM Conference on Hypertext and Social Media (Barcelona, Spain) (HT ’22). Assoc...
2022
-
[23]
Gabriel Magno and Virgilio Almeida. 2021. Measuring international online human values with word embeddings. ACM Transactions on the Web (TWEB) 16, 2 (2021), 1–38
2021
-
[24]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781 [cs.CL] https://arxiv.org/abs/1301.3781
2013 arXiv
-
[25]
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed Representations of Words and Phrases and their Compositional- ity. In Advances in Neural Information Processing Systems , C.J. Burges, L. Bot- tou, M. Welling, Z. Ghahramani, and K.Q. Wei...
2013
-
[26]
Harvard Institute of Politics. 2024. Harvard Youth Poll. https://iop.harvard.edu/ youth-poll/48th-edition-fall-2024 Accessed: 2024-11-20
2024
-
[27]
Raphael Ottoni, Evandro Cunha, Gabriel Magno, Pedro Bernardina, Wagner Meira Jr., and Virgílio Almeida. 2018. Analyzing Right-wing YouTube Channels: Hate, Violence and Discrimination. In Proceedings of the 10th ACM Conference on Web Science (Amsterdam, Netherlands) (WebSci ’18...
2018 doi
-
[28]
Raphael Ottoni, Evandro Cunha, Gabriel Magno, Pedro Bernardina, Wagner Meira Jr, and Virgílio Almeida. 2018. Analyzing right-wing youtube channels: Hate, violence and discrimination. In Proceedings of the 10th ACM conference on Analyzing Political Discourse on Discord during t...
2018
-
[29]
Mark O’Brien. 2023. The coming of the storm: moral panics, social media and regulation in the QAnon era. Information & Communications Technology Law 32, 1 (2023), 102–121
2023
-
[30]
Kevin Roose. 2017. This Was the Alt-Right’s Favorite Chat App. Then Came Charlottesville. https://www.nytimes.com/2017/08/15/technology/discord-chat- app-alt-right.html
2017
-
[31]
Punyajoy Saha, Mithun Das, Binny Mathew, and Animesh Mukherjee. 2023. Hate Speech: Detection, Mitigation and Beyond. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining (Singapore, Singa- pore) (WSDM ’23). Association for Computing Machin...
2023
-
[32]
Ashwini Kumar Singh, Vahid Ghafouri, Jose Such, Guillermo Suarez-Tangil, et al
-
[33]
Robyn Speer, Joshua Chin, and Catherine Havasi. 2017. ConceptNet 5.5: An Open Multilingual Graph of General Knowledge. 4444–4451. http://aaai.org/ocs/index. php/AAAI/AAAI17/paper/view/14972
2017
-
[34]
Peter Stefanov, Kareem Darwish, Atanas Atanasov, and Preslav Nakov. 2020. Predicting the Topical Stance and Political Leaning of Media using Tweets. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Sc...
2020
-
[35]
The Editorial Board. 2024. Opinion | The Only Patriotic Choice for President. The New York Times (2024). https://www.nytimes.com/2024/09/30/opinion/editorials/ kamala-harris-2024-endorsement.html
2024
-
[36]
Richard Ashby Wilson and Molly K Land. 2020. Hate speech on social media: Content moderation in context. Conn. L. Rev. 52 (2020), 1029
2020
-
[37]
Gadi Wolfsfeld, Elad Segev, and Tamir Sheafer. 2013. Social media and the Arab Spring: Politics comes first. The International Journal of Press/Politics 18, 2 (2013), 115–137
2013
-
[38]
I Won the Election!
Savvas Zannettou. 2021. " I Won the Election!": an empirical analysis of soft moderation interventions on Twitter. In Proceedings of the international AAAI conference on web and social media , Vol. 15. 865–876
2021
-
[2021]
Institute for Strategic Dialogue 19 (2021)
The extreme right on discord. Institute for Strategic Dialogue 19 (2021)
2021
-
[2024]
In International AAAI Conference on Web and Social Media
Differences in the Toxic Language of Cross-Platform Communities. In International AAAI Conference on Web and Social Media
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.