Pith. sign in

REVIEW 3 major objections 5 minor 160 references

A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Survey maps the state of hate speech detection in low-resource languages, cataloging datasets and methods.

desk verdict Useful updated map of datasets and techniques for low-resource hate speech detection, but the coverage claim is unverifiable as reported and the scope is overbroad. read the letter →

arxiv 2411.19017 v1 pith:N2OXGVSJ submitted 2024-11-28 cs.CL cs.LG

classification cs.CLcs.LG
keywords hatespeechdetectionlow-resourcelanguagesdatasetsurveynaturallanguageprocessingdeeplearningmultilingualIndicsocialmedia
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a survey rather than a new detection method. Its goal is to establish a current, structured picture of automatic hate speech detection for low-resource languages: which datasets exist, which features and machine- or deep-learning techniques have been tried, and where the research gaps are. The authors organize the literature by continent and language family, give special attention to Indic languages, and draw a trajectory from TF-IDF and classical classifiers toward word embeddings, LSTM variants, and transformer models. If the survey is faithful, it gives researchers an entry point that previously had to be assembled from many separate papers.

What carries the argument

The load-bearing device is the survey's taxonomy, built through a keyword-driven search and manual filtering. Datasets are sorted into English, monolingual low-resource, multilingual, and multimodal buckets; detection studies are sorted by continent and, for Asia, into Indic and non-Indic languages. This structure lets the paper convert a collection of individual papers into comparative claims about resource availability, dominant methods, and open problems.

What would settle it

Replicating the Section 3.1 search with the listed keywords in the named digital libraries and checking whether the catalog omits a substantial body of published low-resource-language datasets or detection studies, particularly from regional venues, would settle whether the survey's coverage claim holds.

Watch

Extended reading notes

Core claim

The paper's central claim is that hate speech detection research is overwhelmingly English-centric and that the work on low-resource languages, though growing, is scattered and uneven. It presents a catalog of monolingual, multilingual, and multimodal hate speech datasets from European, Latin American, African, and Asian languages, together with the feature representations and models applied to them. It also claims that the field has moved from TF-IDF and conventional machine learning to embeddings combined with deep models, with multilingual transformers such as BERT variants and XLM-RoBERTa as the current standard. The survey's accompanying thesis is that the main obstacles are not only scarce data but also culturally specific meaning, ambiguous annotation, biased collection, and the near absence of public code.

Load-bearing premise

The survey's usefulness depends on its literature search having found a representative sample of the work on hate speech detection in low-resource languages, since the exact queries, databases, and screening rules are only partially documented in Section 3.1.

Editorial extensions

If this is right

  • Researchers working on a low-resource language can use the dataset tables to locate existing corpora and shared-task benchmarks before building new ones.
  • The reported shift from TF-IDF and classical machine learning to word embeddings and fine-tuned transformers implies that new work should start with pre-trained multilingual or monolingual models.
  • For Indic languages, shared tasks such as HASOC provide reusable evaluation infrastructure, so progress can be measured against common baselines.
  • The challenges the survey lists imply that better detection requires culturally informed annotation and dataset documentation, not just larger models.
  • Multimodal hate speech, especially memes and code-mixed text, is an open area where the paper predicts further growth.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The continent-wise organization makes explicit that 'low-resource' covers very unequal situations: some European languages resemble English in data availability, while many African and Southeast Asian languages have only a few small datasets; a resource-ranking study could quantify this.
  • The paper's emphasis on code-mixed and transliterated Indic text suggests a testable extension: models trained on code-mixed Roman-script data may transfer poorly to native-script text unless both scripts appear in training.
  • Because the survey relies on English-indexed digital libraries, its coverage likely underrepresents work published in regional venues or local languages; searching regional databases would test this gap.
  • The catalog could be turned into a living benchmark: standardizing annotation labels across datasets would let researchers compare cross-lingual transfer results more reliably.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper is a survey of automatic online hate speech detection, with a stated focus on low-resource languages. It reviews definitions of hate speech from major platforms, proposes a categorization with overlapping concepts, describes a keyword-based literature search, catalogs English and non-English datasets (monolingual, multilingual, and multimodal), summarizes detection methods by region (Europe, Latin America, Africa, Asia, and the Indian subcontinent), and closes with research challenges and future directions. The central claim, stated in the abstract, is that the article provides a detailed and current survey of hate speech detection in low-resource languages, including available datasets, features, and techniques.

Significance. If the coverage claims are trustworthy, the survey would be a useful entry point for researchers working on non-English hate speech detection, especially because of its consolidated dataset tables (Tables 4 and 5), its continent-wise organization, and its attention to Indic languages. The paper's main strengths are breadth of languages surveyed and the structured tabulation of many datasets and methods. However, the survey's core value depends on two conditions: that the literature search is systematic enough to support the word 'detailed'/'comprehensive,' and that the term 'low-resource' is applied consistently. The manuscript currently does not meet either condition: the search protocol in Section 3.1 is under-specified, and Section 5.2 explicitly classifies Spanish, French, German, Portuguese, and other relatively well-resourced languages as low-resource. There is also a concrete fidelity error in Table 2. The paper provides no machine-checked proofs, reproducible code, or parameter-free derivations; its contribution is a narrative synthesis of the secondary literature.

major comments (3)
  1. [Section 3.1 (3.1.1 and 3.1.2)] The literature search is not described in enough detail to support the claim of a comprehensive survey. The manuscript lists broad database names and a table of keyword categories, but gives no exact query strings, no search dates, no number of records retrieved per database, no deduplication counts, and no screening counts. The exclusion criterion in Section 3.2, 'Research papers with indeterminate information were excluded,' is not defined. Because the contribution is explicitly a survey, this undocumented pipeline makes the coverage claim unfalsifiable and unreproducible. The authors should add a full search log, inclusion/exclusion criteria, and a flow diagram, or substantially soften the comprehensiveness claim to that of a selective review.
  2. [Section 5.2] The definition of 'low-resource' is internally inconsistent: 'except for English, we have considered all the other languages as low-resource languages' includes Spanish, French, German, Portuguese, Arabic, and Korean, several of which are among the most resourced non-English languages. This contradicts the paper's own language-distribution charts (Fig. 4), where Spanish, German, and Arabic are shown as the second-most-used languages on major platforms. The broadened scope dilutes the stated focus and changes what the survey's conclusions about 'low-resource languages' actually mean. The authors should adopt an operational definition (e.g., based on dataset size, NLP tool coverage, or speaker population) and apply it consistently throughout.
  3. [Table 2, row 9] The example 'People from <insert religion> should not be allowed inside.' is described in the text as religious hate speech that is neither abusive nor cyberbullying, but the table marks Religion as 'No' and Ableism as 'Yes'. This is a direct contradiction between the table and the text. The row must be corrected, and the remaining rows should be re-checked for similar inconsistencies.
minor comments (5)
  1. [Abstract] The abstract ends with 'Keywords: article, template, simple', which appears to be a leftover template string and should be removed before submission.
  2. [Section 4.2.1, Tulkens dataset] The label distribution sums incorrectly: 344 invalid comments out of 6375 is approximately 5.4%, not 0.05% as printed. Please correct the percentage and verify the arithmetic in the surrounding dataset descriptions.
  3. [Section 3.2] The phrase 'Research papers with indeterminate information were excluded' is vague; please specify what information was considered indeterminate and how this determination was made.
  4. [Section 5.2.3, Amharic] The sentence 'The Word2Vec model with Naïve Bayes classifier achieved the best performance with accuracy and ROC Score of 79.83% and 0.8305, respectively' is clear, but the preceding sentence about the Apache Spark-based model would benefit from a citation to the specific model description or the paper being summarized.
  5. [Section 5.2.4, Japanese subsection] The description of Kim et al. [82] refers to 'racist hate speech' measured with 'adoption thresholds of racist hate speech,' but the paper's focus on Korean racism may be clearer if the target group were named explicitly; please check that the summary matches the cited paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey synthesizes external literature and contains no derivation, fitted parameter, or prediction that reduces to its own inputs.

full rationale

This is a literature survey with no formal derivation chain, no fitted parameters, and no predictions generated from a model. Its contribution is a synthesis of externally published datasets, features, and methods, summarized in Tables 4-6 and Sections 4-5, so there is no self-referential fitting loop. The collection procedure in Section 3.1 is a keyword-based search of external databases; even if the search is underspecified or non-reproducible, that is a completeness and verifiability limitation, not circular reasoning. The scope statement in Section 5.2, 'except for English, we have considered all the other languages as low-resource languages,' is a definitional choice rather than a hidden input that forces the survey's conclusions. The Table 2 inconsistency, where a religious hate speech example is marked 'Religion: No' and 'Ableism: Yes,' is an internal data-quality or annotation error, not a reduction of a claimed result to its own assumptions. No load-bearing self-citation was identified: references to 'Das et al.' point to Mithun Das, Amit Kumar Das, and other researchers, not to the present authors, and none of the survey's organizing claims depends on a uniqueness theorem or ansatz imported from the same author group. The paper is therefore not circular.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey's factual content relies on the fidelity of the cited publications and on third-party statistics. These are domain assumptions rather than mathematical axioms; if any major cited summary is misreported, the corresponding portion of the survey is wrong. No ad hoc-to-paper entities are introduced.

assumptions (3)
  • domain assumption The hate speech definitions quoted from Meta, YouTube, X, TikTok, and LinkedIn are representative and accurately transcribed.
    Section 2.1 relies on these definitions to build the categorization in Table 1.
  • domain assumption Public statistics on language usage and hate speech actions (Figs 2, 4, 6, 7) from Statista, Wikipedia, and IPSOS are accurate as cited.
    Figures 2, 4, 6, and 7 are used to motivate the low-resource language focus, but the underlying data are third-party and not independently verified.
  • domain assumption The keyword-based search strategy in Section 3.1 yields a representative sample of the hate speech detection literature in low-resource languages.
    The survey's comprehensiveness claim depends on the completeness of the search; the exact queries, database coverage, and exclusion criteria are only partially specified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages." pith.science (2026). https://pith.science/paper/N2OXGVSJ

@misc{pith2026241119017,
  author       = {Pith},
  title        = {Pith review of: A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N2OXGVSJ}},
  note         = {Machine review of arXiv:2411.19017}
}
read the original abstract

The expanding influence of social media platforms over the past decade has impacted the way people communicate. The level of obscurity provided by social media and easy accessibility of the internet has facilitated the spread of hate speech. The terms and expressions related to hate speech gets updated with changing times which poses an obstacle to policy-makers and researchers in case of hate speech identification. With growing number of individuals using their native languages to communicate with each other, hate speech in these low-resource languages are also growing. Although, there is awareness about the English-related approaches, much attention have not been provided to these low-resource languages due to lack of datasets and online available data. This article provides a detailed survey of hate speech detection in low-resource languages around the world with details of available datasets, features utilized and techniques used. This survey further discusses the prevailing surveys, overlapping concepts related to hate speech, research challenges and opportunities.

Figures

Figures reproduced from arXiv: 2411.19017 by the authors.

Figure 1
Figure 1. Survey Overview The principal objective of this survey is to provide a comprehensive idea about the history of research on hate speech and a detailed account of its application in case of low-resource languages. This survey paper is organized as follows: Section 2 consists background study detailing the definition of hate speech as explained by different social media giants, categorizing hate speech and the overlapp… view at source ↗
Figure 2
Figure 2. Monthly Active Users on Different Social Media Platforms[42] [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Relation between Hate Speech and Extended Concepts [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Content Percentage of Languages on Various Social Media Platforms [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Collection of Relevant Documents 3.3 Literature Review of Survey/Review Articles Surveys and review articles are eminent means of gaining knowledge and resources regarding a particular research field. We collected review articles regarding hate speech and abusive langu…
Figure 6
Figure 6. Figure 6: Action Taken Against Hate Posts on Meta Platforms[38][41] [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Hate Speech on Social Media across 16 countries[63] [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

160 extracted references · 60 canonical work pages

  1. [1]

    Problematic content in spanish language comments in youtube videos about venezuelan refugees and migrants.Journal of Quantitative Description: Digital Media , 1, 2021

    Luis Aguirre and Emese Domahidi. Problematic content in spanish language comments in youtube videos about venezuelan refugees and migrants.Journal of Quantitative Description: Digital Media , 1, 2021

  2. [2]

    Detecting hate speech against women in english tweets

    Resham Ahluwalia, Himani Soni, Edward Callow, Anderson Nascimento, and Martine De Cock. Detecting hate speech against women in english tweets. EV ALITA Evaluation of NLP and Speech Tools for Italian, 12:194, 2018

  3. [3]

    Automatic detection of offensive language for urdu and roman urdu

    Muhammad Pervez Akhter, Zheng Jiangbin, Irfan Raza Naqvi, Mohammed Abdelmajeed, and Muhammad Tariq Sadiq. Automatic detection of offensive language for urdu and roman urdu. IEEE Access, 8:91213–91226, 2020

  4. [4]

    Detection of hate speech in social networks: a survey on multilingual corpus

    Areej Al-Hassan and Hmood Al-Dossari. Detection of hate speech in social networks: a survey on multilingual corpus. In 6th international conference on computer science and information technol- ogy, volume 10, pages 10–5121. ACM, 2019

  5. [5]

    Dataset construction for the detection of anti-social behaviour in online communication in arabic

    Azalden Alakrot, Liam Murray, and Nikola S Nikolov. Dataset construction for the detection of anti-social behaviour in online communication in arabic. Procedia Computer Science, 142:174–181, 2018

  6. [6]

    Are they our brothers? analysis and detection of religious hate speech in the arabic twittersphere

    Nuha Albadi, Maram Kurdi, and Shivakant Mishra. Are they our brothers? analysis and detection of religious hate speech in the arabic twittersphere. In 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM) , pages 69–76. IEEE, 2018

  7. [7]

    Edwin Aldana-Bobadilla, Alejandro Molina-Villegas, Yuridia Montelongo-Padilla, Ivan Lopez- Arevalo, and Oscar S. Sordia. A language model for misogyny detection in latin american spanish driven by multisource feature extraction and transformers. Applied Sciences, 11(21):10467, 2021

  8. [8]

    Hate speech detection in the indonesian language: A dataset and preliminary study

    Ika Alfina, Rio Mulia, Mohamad Ivan Fanany, and Yudo Ekanata. Hate speech detection in the indonesian language: A dataset and preliminary study. In 2017 international conference on advanced computer science and information systems (ICACSIS) , pages 233–238. IEEE, 2017

Show all 160 references
  1. [9]

    Improving hate speech detection of urdu tweets using sentiment analysis

    Muhammad Z Ali, Sahar Rauf, Kashif Javed, Sarmad Hussain, et al. Improving hate speech detection of urdu tweets using sentiment analysis. IEEE Access, 9:84296–84305, 2021

  2. [10]

    A literature review of textual hate speech detection methods and datasets

    Fatimah Alkomah and Xiaogang Ma. A literature review of textual hate speech detection methods and datasets. Information, 13(6):273, 2022

  3. [11]

    Deep learning-based hate speech detection in code-mixed tamil text

    S Anbukkarasi and S Varadhaganapathy. Deep learning-based hate speech detection in code-mixed tamil text. IETE Journal of Research , 69(11):7893–7898, 2023

  4. [12]

    Multi- modal hate speech detection in memes using contrastive language-image pre-training.IEEE Access, 2024

    Greeshma Arya, Mohammad Kamrul Hasan, Ashish Bagwari, Nurhizam Safie, Shayla Islam, Fa- tima Rayan Awad Ahmed, Aaishani De, Muhammad Attique Khan, and Taher M Ghazal. Multi- modal hate speech detection in memes using contrastive language-image pre-training.IEEE Access, 2024

  5. [13]

    Model-agnostic meta-learning for multilingual hate speech detection

    Md Rabiul Awal, Roy Ka-Wei Lee, Eshaan Tanwar, Tanmay Garg, and Tanmoy Chakraborty. Model-agnostic meta-learning for multilingual hate speech detection. IEEE Transactions on Com- putational Social Systems , 11(1):1086–1095, 2023

  6. [14]

    Toxicity detection on bengali social media comments using supervised models

    Nayan Banik and Md Hasan Hafizur Rahman. Toxicity detection on bengali social media comments using supervised models. In 2019 2nd international conference on Innovation in Engineering and Technology (ICIET), pages 1–5. IEEE, 2019

  7. [15]

    Qutnocturnal@ hasoc’19: Cnn for hate speech and offensive content identification in hindi language

    Md Abul Bashar and Richi Nayak. Qutnocturnal@ hasoc’19: Cnn for hate speech and offensive content identification in hindi language. arXiv preprint arXiv:2008.12448 , 2020

  8. [16]

    Semeval-2019 task 5: Multilin- gual detection of hate speech against immigrants and women in twitter

    Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. Semeval-2019 task 5: Multilin- gual detection of hate speech against immigrants and women in twitter. In Proceedings of the 13t...

  9. [17]

    Building a formal model for hate detection in french corpora

    Delphine Battistelli, Cyril Bruneau, and Valentina Dragos. Building a formal model for hate detection in french corpora. Procedia Computer Science, 176:2358–2365, 2020

  10. [18]

    Interpretable multi labeled bengali toxic comments classification using deep learning

    Tanveer Ahmed Belal, GM Shahariar, and Md Hasanul Kabir. Interpretable multi labeled bengali toxic comments classification using deep learning. In 2023 International Conference on Electrical, Computer and Communication Engineering (ECCE) , pages 1–6. IEEE, 2023

  11. [19]

    A turkish hate speech dataset and detection system

    Fatih Beyhan, Buse C ¸ arık, Inan¸ c Arın, Ay¸ secan Terzio˘ glu, Berrin Yanikoglu, and Reyyan Yeniterzi. A turkish hate speech dataset and detection system. In Proceedings of the thirteenth language resources and evaluation conference, pages 4177–4185, 2022

  12. [20]

    One to rule them all: Towards joint indic language hate speech detection

    Mehar Bhatia, Tenzin Singhay Bhotia, Akshat Agarwal, Prakash Ramesh, Shubham Gupta, Kumar Shridhar, Felix Laumann, and Ayushman Dash. One to rule them all: Towards joint indic language hate speech detection. arXiv preprint arXiv:2109.13711 , 2021

  13. [21]

    A dataset of hindi-english code-mixed social media text for hate speech detection

    Aditya Bohra, Deepanshu Vijay, Vinay Singh, Syed Sarfaraz Akhtar, and Manish Shrivastava. A dataset of hindi-english code-mixed social media text for hate speech detection. In Proceedings of the second workshop on computational modeling of people’s opinions, personality, and e...

  14. [22]

    Hatebert: Retraining bert for abusive language detection in english

    Tommaso Caselli, Valerio Basile, Jelena Mitrovi´ c, and Michael Granitzer. Hatebert: Retraining bert for abusive language detection in english. arXiv preprint arXiv:2010.12472 , 2020

  15. [23]

    Analyzing zero-shot transfer scenarios across spanish variants for hate speech detection

    Galo Castillo-L´ opez, Arij Riabi, and Djam´ e Seddah. Analyzing zero-shot transfer scenarios across spanish variants for hate speech detection. In Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023) , pages 1–13, 2023

  16. [24]

    Tiktok: languages covered by content moderators, 2024

    Laura Ceci. Tiktok: languages covered by content moderators, 2024. URL https://www. statista.com/statistics/1405643/tiktok-language-covered-moderators/ . Accessed: 2024- 08-28

  17. [25]

    Policycorpus xl: An italian corpus for the detection of hate speech against politics

    Fabio Celli, Mirko Lai, Armend Duzha, Cristina Bosco, Viviana Patti, et al. Policycorpus xl: An italian corpus for the detection of hate speech against politics. In CEUR workshop proceedings, volume 3033, pages 1–7. CEUR-WS. org, 2021

  18. [26]

    Offensive language identification in dravidian languages using mpnet and cnn

    Bharathi Raja Chakravarthi, Manoj Balaji Jagadeeshan, Vasanth Palanikumar, and Ruba Priyad- harshini. Offensive language identification in dravidian languages using mpnet and cnn. Interna- tional Journal of Information Management Data Insights , 3(1):100151, 2023

  19. [27]

    A twitter bert approach for offensive language detection in marathi

    Tanmay Chavan, Shantanu Patankar, Aditya Kane, Omkar Gokhale, and Raviraj Joshi. A twitter bert approach for offensive language detection in marathi. arXiv preprint arXiv:2212.10039 , 2022

  20. [28]

    A literature survey on multimodal and multi- lingual automatic hate speech identification

    Anusha Chhabra and Dinesh Kumar Vishwakarma. A literature survey on multimodal and multi- lingual automatic hate speech identification. Multimedia Systems, 29(3):1203–1230, 2023

  21. [29]

    Conan–counter narratives through nichesourcing: a multilingual dataset of responses to fight online hate speech

    Yi-Ling Chung, Elizaveta Kuzmenko, Serra Sinem Tekiroglu, and Marco Guerini. Conan–counter narratives through nichesourcing: a multilingual dataset of responses to fight online hate speech. arXiv preprint arXiv:1910.03270 , 2019

  22. [30]

    A corpus of turkish offensive language on social media

    C ¸ a˘ grı C ¸ ¨ oltekin. A corpus of turkish offensive language on social media. InProceedings of the twelfth language resources and evaluation conference , pages 6174–6184, 2020

  23. [31]

    Bangla hate speech detection on social media using attention-based recurrent neural network

    Amit Kumar Das, Abdullah Al Asif, Anik Paul, and Md Nur Hossain. Bangla hate speech detection on social media using attention-based recurrent neural network. Journal of Intelligent Systems , 30 (1):578–591, 2021

  24. [32]

    Data bootstrapping approaches to improve low resource abusive language detection for indic languages

    Mithun Das, Somnath Banerjee, and Animesh Mukherjee. Data bootstrapping approaches to improve low resource abusive language detection for indic languages. In Proceedings of the 33rd ACM conference on hypertext and social media , pages 32–42, 2022

  25. [33]

    Hate speech and offen- sive language detection in bengali

    Mithun Das, Somnath Banerjee, Punyajoy Saha, and Animesh Mukherjee. Hate speech and offen- sive language detection in bengali. arXiv preprint arXiv:2210.03479 , 2022. 26 This work is shared under a CC BY-SA 4.0 license unless otherwise noted

  26. [34]

    Low-resource counterspeech generation for indic languages: The case of bengali and hindi

    Mithun Das, Saurabh Kumar Pandey, Shivansh Sethi, Punyajoy Saha, and Animesh Mukherjee. Low-resource counterspeech generation for indic languages: The case of bengali and hindi. arXiv preprint arXiv:2402.07262, 2024

  27. [35]

    Automated hate speech de- tection and the problem of offensive language

    Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. Automated hate speech de- tection and the problem of offensive language. In Proceedings of the international AAAI conference on web and social media , volume 11, pages 512–515, 2017

  28. [36]

    Sentiment analysis methods for politics and hate speech contents in spanish language: a systematic review

    Ernesto del Valle and Luis de la Fuente. Sentiment analysis methods for politics and hate speech contents in spanish language: a systematic review. IEEE Latin America Transactions , 21(3): 408–418, 2023

  29. [37]

    Detox: A comprehensive dataset for german offensive language and conversation analysis

    Christoph Demus, Jonas Pitz, Mina Sch¨ utz, Nadine Probol, Melanie Siegel, and Dirk Labudde. Detox: A comprehensive dataset for german offensive language and conversation analysis. In Proceedings of the Sixth Workshop on Online Abuse and Harms (WOAH) , pages 143–153, 2022

  30. [38]

    Statista dataset for facebook, 2023

    Statista Research Department. Statista dataset for facebook, 2023. URL https://www.statista. com/statistics/1013804/facebook-hate-speech-content-deletion-quarter/ . Accessed: 2024-08-28

  31. [39]

    Hate speech detection in asian languages: a survey

    LK Dhanya and Kannan Balakrishnan. Hate speech detection in asian languages: a survey. In 2021 international conference on communication, control and information sciences (ICCISc) , volume 1, pages 1–5. IEEE, 2021

  32. [40]

    Com- mon sense reasoning for detection, prevention, and mitigation of cyberbullying

    Karthik Dinakar, Birago Jones, Catherine Havasi, Henry Lieberman, and Rosalind Picard. Com- mon sense reasoning for detection, prevention, and mitigation of cyberbullying. ACM Transactions on Interactive Intelligent Systems (TiiS) , 2(3):1–30, 2012

  33. [41]

    Statista dataset for instagram, 2023

    Stacy Jo Dixon. Statista dataset for instagram, 2023. URL https://www.statista.com/ statistics/1275933/global-actioned-hate-speech-content-instagram/ . Accessed: 2024- 08-28

  34. [42]

    Statista most used social networks, 2024

    Stacy Jo Dixon. Statista most used social networks, 2024. URL https://www.statista. com/statistics/272014/global-social-networks-ranked-by-number-of-users/ . Accessed: 2024-08-28

  35. [43]

    Hate speech detection on vietnamese social media text using the bidirectional-lstm model

    Hang Thi-Thuy Do, Huy Duc Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen, and Anh Gia- Tuan Nguyen. Hate speech detection on vietnamese social media text using the bidirectional-lstm model. arXiv preprint arXiv:1911.03648 , 2019

  36. [44]

    At the lower end of language—exploring the vulgar and obscene side of german

    Elisabeth Eder, Ulrike Krieg-Holz, and Udo Hahn. At the lower end of language—exploring the vulgar and obscene side of german. In Proceedings of the third workshop on abusive language online, pages 119–128, 2019

  37. [45]

    A deep learning approach to detect abusive bengali text

    Estiak Ahmed Emon, Shihab Rahman, Joti Banarjee, Amit Kumar Das, and Tanni Mittra. A deep learning approach to detect abusive bengali text. In 2019 7th International Conference on Smart Computing & Communications (ICSCC) , pages 1–5. IEEE, 2019

  38. [46]

    A study on the feasibility to detect hate speech in swedish

    Johan Fernquist, Oskar Lindholm, Lisa Kaati, and Nazar Akrami. A study on the feasibility to detect hate speech in swedish. In 2019 IEEE international conference on big data (Big Data) , pages 4724–4729. IEEE, 2019

  39. [47]

    A survey on automatic detection of hate speech in text

    Paula Fortuna and S´ ergio Nunes. A survey on automatic detection of hate speech in text. ACM Computing Surveys (CSUR) , 51(4):1–30, 2018

  40. [48]

    A hierarchically-labeled portuguese hate speech dataset

    Paula Fortuna, Joao Rocha da Silva, Leo Wanner, S´ ergio Nunes, et al. A hierarchically-labeled portuguese hate speech dataset. In Proceedings of the third workshop on abusive language online , pages 94–104, 2019

  41. [49]

    Defining disability: Understandings of and attitudes towards ableism and disability

    Carli Friedman and Aleksa L Owen. Defining disability: Understandings of and attitudes towards ableism and disability. Disability Studies Quarterly , 37(1), 2017

  42. [50]

    Normalizing misogyny: hate speech and verbal abuse of female politicians on japanese twitter

    Tamara Fuchs and Fabian Sch¨ afer. Normalizing misogyny: hate speech and verbal abuse of female politicians on japanese twitter. In Japan forum, volume 33, pages 553–579. Taylor & Francis, 2021. 27 This work is shared under a CC BY-SA 4.0 license unless otherwise noted

  43. [51]

    Cross- lingual offensive language identification for low resource languages: The case of marathi

    Saurabh Gaikwad, Tharindu Ranasinghe, Marcos Zampieri, and Christopher M Homan. Cross- lingual offensive language identification for low resource languages: The case of marathi. arXiv preprint arXiv:2109.03552, 2021

  44. [52]

    Hate speech detection: a comparison of mono and multi- lingual transformer model with cross-language evaluation

    Koyel Ghosh and Apurbalal Senapati. Hate speech detection: a comparison of mono and multi- lingual transformer model with cross-language evaluation. In Proceedings of the 36th Pacific Asia Conference on Language, Information and Computation , pages 853–865, 2022

  45. [53]

    Annihilate hates (task 4 hasoc 2023): Hate speech detection in assamese bengali and bodo languages

    Koyel Ghosh, Apurbalal Senapati, and Aditya Shankar Pal. Annihilate hates (task 4 hasoc 2023): Hate speech detection in assamese bengali and bodo languages. In FIRE (Working Notes) , pages 368–382, 2023

  46. [54]

    Transformer-based hate speech detection in assamese

    Koyel Ghosh, Debarshi Sonowal, Abhilash Basumatary, Bidisha Gogoi, and Apurbalal Senapati. Transformer-based hate speech detection in assamese. In 2023 IEEE Guwahati Subsection Confer- ence (GCON), pages 1–5. IEEE, 2023

  47. [55]

    Sehc: A benchmark setup to identify online hate speech in english

    Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya, Tista Saha, Alka Kumar, and Shikha Sri- vastava. Sehc: A benchmark setup to identify online hate speech in english. IEEE Transactions on Computational Social Systems , 10(2):760–770, 2022

  48. [56]

    Exploring hate speech detection in multimodal publications

    Raul Gomez, Jaume Gibert, Lluis Gomez, and Dimosthenis Karatzas. Exploring hate speech detection in multimodal publications. In Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages 1470–1478, 2020

  49. [57]

    Multilingual abusive comment detection at scale for indic languages

    Vikram Gupta, Sumegh Roychowdhury, Mithun Das, Somnath Banerjee, Punyajoy Saha, Binny Mathew, Animesh Mukherjee, et al. Multilingual abusive comment detection at scale for indic languages. Advances in Neural Information Processing Systems , 35:26176–26191, 2022

  50. [58]

    Kancmd: Kannada codemixed dataset for sentiment analysis and offensive language detection

    Adeep Hande, Ruba Priyadharshini, and Bharathi Raja Chakravarthi. Kancmd: Kannada codemixed dataset for sentiment analysis and offensive language detection. In Proceedings of the Third Workshop on Computational Modeling of People’s Opinions, Personality, and Emotion ’s in Soci...

  51. [59]

    Multi-class sentiment classification on bengali social media comments using machine learning

    Rezaul Haque, Naimul Islam, Mayisha Tasneem, and Amit Kumar Das. Multi-class sentiment classification on bengali social media comments using machine learning. International journal of cognitive computing in engineering , 4:21–35, 2023

  52. [60]

    Multi-class hate speech detection in the norwegian language using fast-rnn and multilingual fine-tuned transformers

    Ehtesham Hashmi and Sule Yildirim Yayilgan. Multi-class hate speech detection in the norwegian language using fast-rnn and multilingual fine-tuned transformers. Complex & Intelligent Systems , 10(3):4535–4556, 2024

  53. [61]

    Mute: A multimodal dataset for detecting hateful memes

    Eftekhar Hossain, Omar Sharif, and Mohammed Moshiul Hoque. Mute: A multimodal dataset for detecting hateful memes. In Proceedings of the 2nd conference of the asia-pacific chapter of the association for computational linguistics and the 12th international joint conference on n...

  54. [62]

    Baad: A multipurpose dataset for automatic bangla offensive speech recognition

    Md Fahad Hossain, Md Al Abid Supto, Zannat Chowdhury, Hana Sultan Chowdhury, and Sheikh Abujar. Baad: A multipurpose dataset for automatic bangla offensive speech recognition. Data in Brief, 48:109067, 2023

  55. [63]

    Survey on the impact of online disinformation and hate speech, 2023

    IPSOS. Survey on the impact of online disinformation and hate speech, 2023. URL https://www.ipsos.com/sites/default/files/ct/news/documents/2023-11/ unesco-ipsos-online-disinformation-hate-speech.pdf . Accessed: 2024-08-28

  56. [64]

    Hateful speech detection in public facebook pages for the bengali language

    Alvi Md Ishmam and Sadia Sharmin. Hateful speech detection in public facebook pages for the bengali language. In 2019 18th IEEE international conference on machine learning and applications (ICMLA), pages 555–560. IEEE, 2019

  57. [65]

    A systematic review of hate speech automatic detection using natural language processing

    Md Saroar Jahan and Mourad Oussalah. A systematic review of hate speech automatic detection using natural language processing. Neurocomputing, page 126232, 2023

  58. [66]

    Banglahatebert: Bert for abusive language detection in bengali

    Md Saroar Jahan, Mainul Haque, Nabil Arhab, and Mourad Oussalah. Banglahatebert: Bert for abusive language detection in bengali. In Proceedings of the second international workshop on resources and techniques for user information in abusive language analysis , pages 8–15, 2022...

  59. [67]

    Finnish hate-speech detection on social media using cnn and finbert

    Md Saroar Jahan, Mourad Oussalah, and Nabil Arhab. Finnish hate-speech detection on social media using cnn and finbert. In Language Resources and Evaluation Conference, LREC 2022, 20-25 June 2022, Palais du Pharo, Marseille, France: conference proceedings . European Language R...

  60. [68]

    Right-wing german hate speech on twitter: Analysis and automatic detection

    Sylvia Jaki and Tom De Smedt. Right-wing german hate speech on twitter: Analysis and automatic detection. arXiv preprint arXiv:1910.07518 , 2019

  61. [69]

    Toxicity detection for indic multilingual social media content

    Manan Jhaveri, Devanshu Ramaiya, and Harveen Singh Chadha. Toxicity detection for indic multilingual social media content. arXiv preprint arXiv:2201.00598 , 2022

  62. [70]

    Harnessing pre-trained sentence transformers for offensive lan- guage detection in indian languages

    Ananya Joshi and Raviraj Joshi. Harnessing pre-trained sentence transformers for offensive lan- guage detection in indian languages. arXiv preprint arXiv:2310.02249 , 2023

  63. [71]

    Bangla hate speech de- tection in videos using machine learning

    Mohd Istiaq Hossain Junaid, Faisal Hossain, and Rashedur M Rahman. Bangla hate speech de- tection in videos using machine learning. In 2021 IEEE 12th Annual Ubiquitous Computing, Elec- tronics & Mobile Communication Conference (UEMCON) , pages 0347–0351. IEEE, 2021

  64. [72]

    Korean online hate speech dataset for multilabel classification: How can social science improve dataset on hate speech? arXiv preprint arXiv:2204.03262 , 2022

    TaeYoung Kang, Eunrang Kwon, Junbum Lee, Youngeun Nam, Junmo Song, and JeongKyu Suh. Korean online hate speech dataset for multilabel classification: How can social science improve dataset on hate speech? arXiv preprint arXiv:2204.03262 , 2022

  65. [73]

    Hatespeech and offensive content detection in hindi language using c-bigru

    Sudharsana Kannan and Jelena Mitrovic. Hatespeech and offensive content detection in hindi language using c-bigru. In FIRE (Working Notes) , pages 209–216, 2021

  66. [74]

    Classifica- tion benchmarks for under-resourced bengali language based on multichannel convolutional-lstm network

    Md Rezaul Karim, Bharathi Raja Chakravarthi, John P McCrae, and Michael Cochez. Classifica- tion benchmarks for under-resourced bengali language based on multichannel convolutional-lstm network. In 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (...

  67. [75]

    Deephateexplainer: Explainable hate speech detection in under-resourced bengali language

    Md Rezaul Karim, Sumon Kanti Dey, Tanhim Islam, Sagor Sarker, Mehadi Hasan Menon, Kabir Hossain, Md Azam Hossain, and Stefan Decker. Deephateexplainer: Explainable hate speech detection in under-resourced bengali language. In 2021 IEEE 8th international conference on data scie...

  68. [76]

    Multimodal hate speech detection from bengali memes and texts

    Md Rezaul Karim, Sumon Kanti Dey, Tanhim Islam, Md Shajalal, and Bharathi Raja Chakravarthi. Multimodal hate speech detection from bengali memes and texts. In International Conference on Speech and Language Technologies for Low-resource Languages , pages 293–308. Springer, 2022

  69. [77]

    G-bert: an efficient method for identifying hate speech in bengali texts on social media

    Ashfia Jannat Keya, Md Mohsin Kabir, Nusrat Jahan Shammey, MF Mridha, Md Rashedul Islam, and Yutaka Watanobe. G-bert: an efficient method for identifying hate speech in bengali texts on social media. IEEE Access, 2023

  70. [78]

    Offensive language detection for low resource language using deep sequence model

    Anas Ali Khan, M Hammad Iqbal, Shibli Nisar, Awais Ahmad, and Waseem Iqbal. Offensive language detection for low resource language using deep sequence model. IEEE Transactions on Computational Social Systems , 2023

  71. [79]

    Muril: Multilin- gual representations for indian languages

    Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, et al. Muril: Multilin- gual representations for indian languages. arXiv preprint arXiv:2103.10730 , 2021

  72. [80]

    Animojity: Detecting hate comments in indic languages and analysing bias against content creators

    Rahul Khurana, Chaitanya Pandey, Priyanshi Gupta, and Preeti Nagrath. Animojity: Detecting hate comments in indic languages and analysing bias against content creators. In Proceedings of the 19th International Conference on Natural Language Processing (ICON) , pages 172–182, 2022

  73. [81]

    The hateful memes challenge: Detecting hate speech in multimodal memes

    Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ring- shia, and Davide Testuggine. The hateful memes challenge: Detecting hate speech in multimodal memes. Advances in neural information processing systems , 33:2611–2624, 2020

  74. [82]

    The impact of politicians’ behaviors on hate speech spread: hate speech adoption threshold on twitter in japan

    Taehee Kim and Yuki Ogawa. The impact of politicians’ behaviors on hate speech spread: hate speech adoption threshold on twitter in japan. Journal of Computational Social Science , pages 1–26, 2024. 29 This work is shared under a CC BY-SA 4.0 license unless otherwise noted

  75. [83]

    Exploiting unsupervised pre-training and automated feature engineering for low-resource hate speech detection in polish

    Renard Korzeniowski, Rafa l Rolczy´ nski, Przemys law Sadownik, Tomasz Korbak, and Marcin Mo˙ zejko. Exploiting unsupervised pre-training and automated feature engineering for low-resource hate speech detection in polish. arXiv preprint arXiv:1906.09325 , 2019

  76. [84]

    Hate speech detection in algerian dialect using deep learning

    Dihia Lanasri, Juan Olano, Sifal Klioui, Sin Liang Lee, and Lamia Sekkai. Hate speech detection in algerian dialect using deep learning. arXiv preprint arXiv:2309.11611 , 2023

  77. [85]

    Project hatemeter: helping ngos and social science researchers to analyze and prevent anti-muslim hate speech on social media

    Mario Laurent. Project hatemeter: helping ngos and social science researchers to analyze and prevent anti-muslim hate speech on social media. Procedia Computer Science , 176:2143–2153, 2020

  78. [86]

    K-mhas: A multi-label hate speech detection dataset in korean online news comment

    Jean Lee, Taejun Lim, Heejun Lee, Bogeun Jo, Yangsok Kim, Heegeun Yoon, and Soyeon Caren Han. K-mhas: A multi-label hate speech detection dataset in korean online news comment. arXiv preprint arXiv:2208.10684, 2022

  79. [87]

    Crehate: Cross-cultural re-annotation of english hate speech dataset

    Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, Juho Kim, and Alice Oh. Crehate: Cross-cultural re-annotation of english hate speech dataset. arXiv preprint arXiv:2308.16705 , 2023

  80. [88]

    Toxic language detection in social media for brazilian portuguese: New dataset and multilingual analysis

    Joao A Leite, Diego F Silva, Kalina Bontcheva, and Carolina Scarton. Toxic language detection in social media for brazilian portuguese: New dataset and multilingual analysis. arXiv preprint arXiv:2010.04543, 2020

  81. [89]

    A large-scale dataset for hate speech detection on vietnamese social media texts

    Son T Luu, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. A large-scale dataset for hate speech detection on vietnamese social media texts. In Advances and Trends in Artificial Intelligence. Artificial Intelligence Practices: 34th International Conference on Industrial, Engineerin...

  82. [90]

    Cyberbullying detection for low-resource languages and dialects: Review of the state of the art

    Tanjim Mahmud, Michal Ptaszynski, Juuso Eronen, and Fumito Masui. Cyberbullying detection for low-resource languages and dialects: Review of the state of the art. Information Processing & Management, 60(5):103454, 2023

  83. [91]

    Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in indo-european languages

    Thomas Mandl, Sandip Modha, Prasenjit Majumder, Daksh Patel, Mohana Dave, Chintak Man- dlia, and Aditya Patel. Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in indo-european languages. In Proceedings of the 11th annual meeting of th...

  84. [92]

    Tueval at semeval-2019 task 5: Lstm approach to hate speech detection in english and spanish

    Mihai Manolescu, Denise L¨ offlad, Adham Nasser Mohamed Saber, and Masoumeh Moradipour Tari. Tueval at semeval-2019 task 5: Lstm approach to hate speech detection in english and spanish. In Proceedings of the 13th International Workshop on Semantic Evaluation , pages 498– 502, 2019

  85. [93]

    Twitter hate speech detection: a systematic review of methods, taxonomy analysis, challenges, and opportunities

    Zainab Mansur, Nazlia Omar, and Sabrina Tiun. Twitter hate speech detection: a systematic review of methods, taxonomy analysis, challenges, and opportunities. IEEE Access, 11:16226– 16249, 2023

  86. [94]

    An ensemble approach for dutch cross-domain hate speech detection

    Ilia Markov, Ine Gevers, and Walter Daelemans. An ensemble approach for dutch cross-domain hate speech detection. In International conference on applications of natural language to information systems, pages 3–15. Springer, 2022

  87. [95]

    Spread of hate speech in online social media

    Binny Mathew, Ritam Dutt, Pawan Goyal, and Animesh Mukherjee. Spread of hate speech in online social media. In Proceedings of the 10th ACM conference on web science , pages 173–182, 2019

  88. [96]

    Hatexplain: A benchmark dataset for explainable hate speech detection

    Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. Hatexplain: A benchmark dataset for explainable hate speech detection. In Proceedings of the AAAI conference on artificial intelligence , volume 35, pages 14867–14875, 2021

  89. [97]

    Did you offend me? classification of offensive tweets in hinglish language

    Puneet Mathur, Ramit Sawhney, Meghna Ayyar, and Rajiv Shah. Did you offend me? classification of offensive tweets in hinglish language. In Proceedings of the 2nd workshop on abusive language online (AL W2), pages 138–148, 2018. 30 This work is shared under a CC BY-SA 4.0 licen...

  90. [98]

    Overview of the hasoc subtrack at fire 2021: Hate speech and offensive content identification in english and indo-aryan languages and conversational hate speech

    Sandip Modha, Thomas Mandl, Gautam Kishore Shahi, Hiren Madhu, Shrey Satapara, Tharindu Ranasinghe, and Marcos Zampieri. Overview of the hasoc subtrack at fire 2021: Hate speech and offensive content identification in english and indo-aryan languages and conversational hate sp...

  91. [99]

    Ethos: a multi- label hate speech detection dataset

    Ioannis Mollas, Zoe Chrysopoulou, Stamatis Karlos, and Grigorios Tsoumakas. Ethos: a multi- label hate speech detection dataset. Complex & Intelligent Systems , 8(6):4663–4678, 2022

  92. [100]

    Social network hate speech detection for amharic language

    Zewdie Mossie, Jenq-Haur Wang, et al. Social network hate speech detection for amharic language. Computer Science & Information Technology , pages 41–55, 2018

  93. [101]

    Abusive language detection on arabic social media

    Hamdy Mubarak, Kareem Darwish, and Walid Magdy. Abusive language detection on arabic social media. In Proceedings of the first workshop on abusive language online , pages 52–56, 2017

  94. [102]

    Hate speech and offensive content detection in indo-aryan languages: A battle of lstm and transformers

    Nikhil Narayan, Mrutyunjay Biswal, Pramod Goyal, and Abhranta Panigrahi. Hate speech and offensive content detection in indo-aryan languages: A battle of lstm and transformers. arXiv preprint arXiv:2312.05671, 2023

  95. [103]

    Offensive language detection in nepali social media

    Nobal B Niraula, Saurab Dulal, and Diwa Koirala. Offensive language detection in nepali social media. In Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021) , pages 67–75, 2021

  96. [104]

    Tackling hate speech in low-resource languages with context experts

    Daniel Nkemelu, Harshil Shah, Michael Best, and Irfan Essa. Tackling hate speech in low-resource languages with context experts. InProceedings of the 2022 International Conference on Information and Communication Technologies and Development , pages 1–11, 2022

  97. [105]

    Abusive language detection in online user content

    Chikashi Nobata, Joel Tetreault, Achint Thomas, Yashar Mehdad, and Yi Chang. Abusive language detection in online user content. In Proceedings of the 25th international conference on world wide web, pages 145–153, 2016

  98. [106]

    Encyclopedia of gender and society , volume 2

    Jodi O’Brien. Encyclopedia of gender and society , volume 2. Sage, 2009

  99. [107]

    Evaluating machine learning techniques for detecting offensive and hate speech in south african tweets

    Oluwafemi Oriola and Eduan Kotz´ e. Evaluating machine learning techniques for detecting offensive and hate speech in south african tweets. IEEE Access, 8:21496–21509, 2020

  100. [108]

    Multilin- gual and multi-aspect hate speech analysis

    Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang, Yangqiu Song, and Dit-Yan Yeung. Multilin- gual and multi-aspect hate speech analysis. arXiv preprint arXiv:1908.11049 , 2019

  101. [109]

    Austrotox: A dataset for target-based austrian german offensive language detection

    Pia Pachinger, Janis Goldzycher, Anna Maria Planitzer, Wojciech Kusa, Allan Hanbury, and Julia Neidhardt. Austrotox: A dataset for target-based austrian german offensive language detection. arXiv preprint arXiv:2406.08080 , 2024

  102. [110]

    Towards multidomain and multi- lingual abusive language detection: a survey

    Endang Wahyu Pamungkas, Valerio Basile, and Viviana Patti. Towards multidomain and multi- lingual abusive language detection: a survey. Personal and Ubiquitous Computing , 27(1):17–43, 2023

  103. [111]

    Regional language toxic comment classification

    Yashkumar Parikh and Jinan Fiaidhi. Regional language toxic comment classification. Authorea Preprints, 2023

  104. [112]

    Multimodal hate speech detection in greek social media

    Konstantinos Perifanos and Dionysis Goutsos. Multimodal hate speech detection in greek social media. Multimodal Technologies and Interaction, 5(7):34, 2021

  105. [113]

    Offensive language identification in greek

    Zeses Pitenis, Marcos Zampieri, and Tharindu Ranasinghe. Offensive language identification in greek. arXiv preprint arXiv:2003.07459 , 2020

  106. [114]

    Comparing pre-trained language models for spanish hate speech detection.Expert Systems with Applications , 166:114120, 2021

    Flor Miriam Plaza-del Arco, M Dolores Molina-Gonz´ alez, L Alfonso Urena-L´ opez, and M Teresa Mart ´ ın-Valdivia. Comparing pre-trained language models for spanish hate speech detection.Expert Systems with Applications , 166:114120, 2021

  107. [115]

    Hate speech annotation: Analysis of an italian twitter corpus

    Fabio Poletto, Marco Stranisci, Manuela Sanguinetti, Viviana Patti, Cristina Bosco, et al. Hate speech annotation: Analysis of an italian twitter corpus. In Ceur workshop proceedings, volume 2006, pages 1–6. CEUR-WS, 2017. 31 This work is shared under a CC BY-SA 4.0 license un...

  108. [116]

    Results of the poleval 2019 shared task 6: First dataset and open shared task for automatic cyberbullying detection in polish twitter

    Michal Ptaszynski, Agata Pieciukiewicz, and Pawe l Dyba la. Results of the poleval 2019 shared task 6: First dataset and open shared task for automatic cyberbullying detection in polish twitter. 2019

  109. [117]

    A benchmark dataset for learning to intervene in online hate speech

    Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding, and William Yang Wang. A benchmark dataset for learning to intervene in online hate speech. arXiv preprint arXiv:1909.04251 , 2019

  110. [118]

    Sold: Sinhala offensive language dataset

    Tharindu Ranasinghe, Isuri Anuradha, Damith Premasiri, Kanishka Silva, Hansi Hettiarachchi, Lasitha Uyangodage, and Marcos Zampieri. Sold: Sinhala offensive language dataset. Language Resources and Evaluation, pages 1–41, 2024

  111. [119]

    A comparative study of different state-of-the-art hate speech detection methods in hindi-english code-mixed data

    Priya Rani, Shardul Suryawanshi, Koustava Goswami, Bharathi Raja Chakravarthi, Theodorus Fransen, and John Philip McCrae. A comparative study of different state-of-the-art hate speech detection methods in hindi-english code-mixed data. In Proceedings of the second workshop on ...

  112. [120]

    Toxlex bn: A curated dataset of bangla toxic language derived from facebook comment

    Mohammad Mamun Or Rashid. Toxlex bn: A curated dataset of bangla toxic language derived from facebook comment. Data in Brief , 43:108416, 2022

  113. [121]

    Offensive language detection using multi-level classification

    Amir H Razavi, Diana Inkpen, Sasha Uritsky, and Stan Matwin. Offensive language detection using multi-level classification. In Advances in Artificial Intelligence: 23rd Canadian Conference on Artificial Intelligence, Canadian AI 2010, Ottawa, Canada, May 31–June 2, 2010. Proce...

  114. [122]

    Hate speech detection in the bengali language: A dataset and its baseline evaluation

    Nauros Romim, Mosahed Ahmed, Hriteshwar Talukder, and Md Saiful Islam. Hate speech detection in the bengali language: A dataset and its baseline evaluation. In Proceedings of International Joint Conference on Advances in Computational Intelligence: IJCACI 2020 , pages 457–468....

  115. [123]

    Bd-shs: A benchmark dataset for learning to detect online bangla hate speech in different social contexts

    Nauros Romim, Mosahed Ahmed, Md Saiful Islam, Arnab Sen Sharma, Hriteshwar Talukder, and Mohammad Ruhul Amin. Bd-shs: A benchmark dataset for learning to detect online bangla hate speech in different social contexts. arXiv preprint arXiv:2206.00372 , 2022

  116. [124]

    Measuring the reliability of hate speech annotations: The case of the european refugee crisis

    Bj¨ orn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki. Measuring the reliability of hate speech annotations: The case of the european refugee crisis. arXiv preprint arXiv:1701.08118 , 2017

  117. [125]

    Hatecheck: Functional tests for hate speech detection models

    Paul R¨ ottger, Bertram Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet B Pierrehumbert. Hatecheck: Functional tests for hate speech detection models. arXiv preprint arXiv:2012.15606, 2020

  118. [126]

    A framework for hate speech detection using deep convolutional neural network

    Pradeep Kumar Roy, Asis Kumar Tripathy, Tapan Kumar Das, and Xiao-Zhi Gao. A framework for hate speech detection using deep convolutional neural network. IEEE Access, 8:204951–204962, 2020

  119. [127]

    Hate speech and offensive language detection in dravidian languages using deep ensemble framework

    Pradeep Kumar Roy, Snehaan Bhawal, and Chinnaudayar Navaneethakrishnan Subalalitha. Hate speech and offensive language detection in dravidian languages using deep ensemble framework. Computer Speech & Language , 75:101386, 2022

  120. [128]

    Detection of hate speech using bert and hate speech word embedding with deep model

    Hind Saleh, Areej Alhothali, and Kawthar Moria. Detection of hate speech using bert and hate speech word embedding with deep model. Applied Artificial Intelligence , 37(1):2166719, 2023

  121. [129]

    Sinhala hate speech detection in social media using text mining and machine learning

    HMST Sandaruwan, SAS Lorensuhewa, and MAL Kalyani. Sinhala hate speech detection in social media using text mining and machine learning. In 2019 19th International Conference on Advances in ICT for Emerging Regions (ICTer) , volume 250, pages 1–8. IEEE, 2019

  122. [130]

    An italian twitter corpus of hate speech against immigrants

    Manuela Sanguinetti, Fabio Poletto, Cristina Bosco, Viviana Patti, and Marco Stranisci. An italian twitter corpus of hate speech against immigrants. In Proceedings of the eleventh international conference on language resources and evaluation (LREC 2018) , 2018

  123. [131]

    Facebook’s top ten languages, 2018

    Juan Pablo Sans. Facebook’s top ten languages, 2018. URL https://www.linkedin.com/pulse/ facebooks-top-ten-languages-who-using-them-juan-pablo# . Accessed: 2024-08-28. 32 This work is shared under a CC BY-SA 4.0 license unless otherwise noted

  124. [132]

    Identifying vulgarity in bengali social media textual content

    Salim Sazzed. Identifying vulgarity in bengali social media textual content. PeerJ Computer Science, 7:e665, 2021

  125. [133]

    Top languages on twitter-stats, 2024

    Semiocast. Top languages on twitter-stats, 2024. URL https://semiocast.com/ top-languages-on-twitter-stats/# . Accessed: 2024-08-28

  126. [134]

    Thar-targeted hate speech against reli- gion: A high-quality hindi-english code-mixed dataset with the application of deep learning models for automatic detection

    Deepawali Sharma, Aakash Singh, and Vivek Kumar Singh. Thar-targeted hate speech against reli- gion: A high-quality hindi-english code-mixed dataset with the application of deep learning models for automatic detection. ACM Transactions on Asian and Low-Resource Language Inform...

  127. [135]

    Targets and aspects in social media hate speech

    Alexander Shvets, Paula Fortuna, Juan Soler, and Leo Wanner. Targets and aspects in social media hate speech. In Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021), pages 179–190, 2021

  128. [136]

    Offensive language and hate speech detection for danish

    Gudbjartur Ingi Sigurbergsson and Leon Derczynski. Offensive language and hate speech detection for danish. arXiv preprint arXiv:1908.04531 , 2019

  129. [137]

    Cyberbullying: Its nature and impact in secondary school pupils

    Peter K Smith, Jess Mahdavi, Manuel Carvalho, Sonja Fisher, Shanette Russell, and Neil Tippett. Cyberbullying: Its nature and impact in secondary school pupils. Journal of child psychology and psychiatry, 49(4):376–385, 2008

  130. [138]

    Detection of hate speech text in hindi-english code- mixed data

    K Sreelakshmi, B Premjith, and KP Soman. Detection of hate speech text in hindi-english code- mixed data. Procedia Computer Science, 171:737–744, 2020

  131. [139]

    Offensive language detection in tamil youtube comments by adapters and cross-domain knowledge transfer

    Malliga Subramanian, Rahul Ponnusamy, Sean Benhur, Kogilavani Shanmugavadivel, Adhithiya Ganesan, Deepti Ravi, Gowtham Krishnan Shanmugasundaram, Ruba Priyadharshini, and Bharathi Raja Chakravarthi. Offensive language detection in tamil youtube comments by adapters and cross-d...

  132. [140]

    Detection of abusive bengali comments for mixed social media data using machine learning

    Sherin Sultana, Md Omur Faruk Redoy, Jabir Al Nahian, Abu Kaisar Mohammad Masum, and Sheikh Abujar. Detection of abusive bengali comments for mixed social media data using machine learning. 2023

  133. [141]

    Indonesia hate speech detection using deep learning

    Taufic Leonardo Sutejo and Dessi Puji Lestari. Indonesia hate speech detection using deep learning. In 2018 International Conference on Asian Language Processing (IALP), pages 39–43. IEEE, 2018

  134. [142]

    Automated amharic hate speech posts and comments detection model using recurrent neural network

    Surafel Getachew Tesfaye and Kula Kakeba. Automated amharic hate speech posts and comments detection model using recurrent neural network. 2020

  135. [143]

    A multi-modal dataset for hate speech detection on social media: Case-study of russia-ukraine conflict

    Surendrabikram Thapa, Aditya Shah, Farhan Ahmad Jafri, Usman Naseem, and Imran Razzak. A multi-modal dataset for hate speech detection on social media: Case-study of russia-ukraine conflict. In CASE 2022-5th Workshop on Challenges and Applications of Automated Extraction of So...

  136. [144]

    A dictionary-based approach to racism detection in dutch social media

    St´ ephan Tulkens, Lisa Hilte, Elise Lodewyckx, Ben Verhoeven, and Walter Daelemans. A dictionary-based approach to racism detection in dutch social media. arXiv preprint arXiv:1608.08738, 2016

  137. [145]

    Quarantining online hate speech: technical and ethical perspectives

    Stefanie Ullmann and Marcus Tomalin. Quarantining online hate speech: technical and ethical perspectives. Ethics and Information Technology , 22(1):69–80, 2020

  138. [146]

    Guidelines for the fine-grained analysis of cyberbullying

    Cynthia Van Hee, Ben Verhoeven, Els Lefever, Guy De Pauw, V´ eronique Hoste, and Walter Daele- mans. Guidelines for the fine-grained analysis of cyberbullying. 2015

  139. [147]

    Detection of racist language in french tweets

    Natalia Vanetik and Elisheva Mimoun. Detection of racist language in french tweets. Information, 13(7):318, 2022

  140. [148]

    Hatebr: A large expert annotated corpus of brazilian instagram comments for offensive language and hate speech detection.arXiv preprint arXiv:2103.14972, 2021

    Francielle Alves Vargas, Isabelle Carvalho, Fabiana Rodrigues de G´ oes, Fabr ´ ıcio Benevenuto, and Thiago Alexandre Salgueiro Pardo. Hatebr: A large expert annotated corpus of brazilian instagram comments for offensive language and hate speech detection.arXiv preprint arXiv:...

  141. [149]

    Towards offensive language identifica- tion for tamil code-mixed youtube comments and posts

    Charangan Vasantharajan and Uthayasanker Thayasivam. Towards offensive language identifica- tion for tamil code-mixed youtube comments and posts. SN Computer Science , 3(1):94, 2022. 33 This work is shared under a CC BY-SA 4.0 license unless otherwise noted

  142. [150]

    Online multilingual hate speech detection: experimenting with hindi and english social media

    Neeraj Vashistha and Arkaitz Zubiaga. Online multilingual hate speech detection: experimenting with hindi and english social media. Information, 12(1):5, 2020

  143. [151]

    Hate and offensive speech detection in hindi and marathi

    Abhishek Velankar, Hrushikesh Patil, Amol Gore, Shubham Salunke, and Raviraj Joshi. Hate and offensive speech detection in hindi and marathi. arXiv preprint arXiv:2110.12200 , 2021

  144. [152]

    L3cube- mahahate: A tweet-based marathi hate speech detection dataset and bert models

    Abhishek Velankar, Hrushikesh Patil, Amol Gore, Shubham Salunke, and Raviraj Joshi. L3cube- mahahate: A tweet-based marathi hate speech detection dataset and bert models. arXiv preprint arXiv:2203.13778, 2022

  145. [153]

    Mono vs multilingual bert for hate speech detection and text classification: A case study in marathi

    Abhishek Velankar, Hrushikesh Patil, and Raviraj Joshi. Mono vs multilingual bert for hate speech detection and text classification: A case study in marathi. In IAPR Workshop on Artificial Neural Networks in Pattern Recognition , pages 121–128. Springer, 2022

  146. [154]

    Hateful symbols or hateful people? predictive features for hate speech detection on twitter

    Zeerak Waseem and Dirk Hovy. Hateful symbols or hateful people? predictive features for hate speech detection on twitter. In Proceedings of the NAACL student research workshop, pages 88–93, 2016

  147. [155]

    German abusive language dataset with focus on covid-19

    Maximilian Wich, Svenja R¨ ather, and Georg Groh. German abusive language dataset with focus on covid-19. In Proceedings of the 17th Conference on Natural Language Processing (KONVENS 2021), pages 247–252, 2021

  148. [156]

    Overview of the germeval 2018 shared task on the identification of offensive language

    Michael Wiegand, Melanie Siegel, and Josef Ruppenhofer. Overview of the germeval 2018 shared task on the identification of offensive language. 2018

  149. [157]

    Content languages on youtube, 2024

    Wikipedia. Content languages on youtube, 2024. URL https://en.wikipedia.org/wiki/ Languages_used_on_the_Internet. Accessed: 2024-08-28

  150. [158]

    Predicting the type and target of offensive posts in social media

    Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. Predicting the type and target of offensive posts in social media. arXiv preprint arXiv:1902.09666, 2019

  151. [159]

    Semeval-2020 task 12: Multilingual offensive language identification in social media (offenseval 2020).arXiv preprint arXiv:2006.07235, 2020

    Marcos Zampieri, Preslav Nakov, Sara Rosenthal, Pepa Atanasova, Georgi Karadzhov, Hamdy Mubarak, Leon Derczynski, Zeses Pitenis, and C ¸ a˘ grı C ¸ ¨ oltekin. Semeval-2020 task 12: Multilingual offensive language identification in social media (offenseval 2020).arXiv preprint ...

  152. [160]

    Detecting hate speech on twitter using a convolution-gru based deep neural network

    Ziqi Zhang, David Robinson, and Jonathan Tepper. Detecting hate speech on twitter using a convolution-gru based deep neural network. In The Semantic Web: 15th International Conference, ESWC 2018, Heraklion, Crete, Greece, June 3–7, 2018, Proceedings 15 , pages 745–760. Springe...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.