REVIEW 3 major objections 5 minor 160 references
A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Survey maps the state of hate speech detection in low-resource languages, cataloging datasets and methods.
desk verdict Useful updated map of datasets and techniques for low-resource hate speech detection, but the coverage claim is unverifiable as reported and the scope is overbroad. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is the survey's taxonomy, built through a keyword-driven search and manual filtering. Datasets are sorted into English, monolingual low-resource, multilingual, and multimodal buckets; detection studies are sorted by continent and, for Asia, into Indic and non-Indic languages. This structure lets the paper convert a collection of individual papers into comparative claims about resource availability, dominant methods, and open problems.
What would settle it
Replicating the Section 3.1 search with the listed keywords in the named digital libraries and checking whether the catalog omits a substantial body of published low-resource-language datasets or detection studies, particularly from regional venues, would settle whether the survey's coverage claim holds.
Extended reading notes
Core claim
The paper's central claim is that hate speech detection research is overwhelmingly English-centric and that the work on low-resource languages, though growing, is scattered and uneven. It presents a catalog of monolingual, multilingual, and multimodal hate speech datasets from European, Latin American, African, and Asian languages, together with the feature representations and models applied to them. It also claims that the field has moved from TF-IDF and conventional machine learning to embeddings combined with deep models, with multilingual transformers such as BERT variants and XLM-RoBERTa as the current standard. The survey's accompanying thesis is that the main obstacles are not only scarce data but also culturally specific meaning, ambiguous annotation, biased collection, and the near absence of public code.
Load-bearing premise
The survey's usefulness depends on its literature search having found a representative sample of the work on hate speech detection in low-resource languages, since the exact queries, databases, and screening rules are only partially documented in Section 3.1.
Editorial extensions
If this is right
- Researchers working on a low-resource language can use the dataset tables to locate existing corpora and shared-task benchmarks before building new ones.
- The reported shift from TF-IDF and classical machine learning to word embeddings and fine-tuned transformers implies that new work should start with pre-trained multilingual or monolingual models.
- For Indic languages, shared tasks such as HASOC provide reusable evaluation infrastructure, so progress can be measured against common baselines.
- The challenges the survey lists imply that better detection requires culturally informed annotation and dataset documentation, not just larger models.
- Multimodal hate speech, especially memes and code-mixed text, is an open area where the paper predicts further growth.
Reading between the lines
- The continent-wise organization makes explicit that 'low-resource' covers very unequal situations: some European languages resemble English in data availability, while many African and Southeast Asian languages have only a few small datasets; a resource-ranking study could quantify this.
- The paper's emphasis on code-mixed and transliterated Indic text suggests a testable extension: models trained on code-mixed Roman-script data may transfer poorly to native-script text unless both scripts appear in training.
- Because the survey relies on English-indexed digital libraries, its coverage likely underrepresents work published in regional venues or local languages; searching regional databases would test this gap.
- The catalog could be turned into a living benchmark: standardizing annotation labels across datasets would let researchers compare cross-lingual transfer results more reliably.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey of automatic online hate speech detection, with a stated focus on low-resource languages. It reviews definitions of hate speech from major platforms, proposes a categorization with overlapping concepts, describes a keyword-based literature search, catalogs English and non-English datasets (monolingual, multilingual, and multimodal), summarizes detection methods by region (Europe, Latin America, Africa, Asia, and the Indian subcontinent), and closes with research challenges and future directions. The central claim, stated in the abstract, is that the article provides a detailed and current survey of hate speech detection in low-resource languages, including available datasets, features, and techniques.
Significance. If the coverage claims are trustworthy, the survey would be a useful entry point for researchers working on non-English hate speech detection, especially because of its consolidated dataset tables (Tables 4 and 5), its continent-wise organization, and its attention to Indic languages. The paper's main strengths are breadth of languages surveyed and the structured tabulation of many datasets and methods. However, the survey's core value depends on two conditions: that the literature search is systematic enough to support the word 'detailed'/'comprehensive,' and that the term 'low-resource' is applied consistently. The manuscript currently does not meet either condition: the search protocol in Section 3.1 is under-specified, and Section 5.2 explicitly classifies Spanish, French, German, Portuguese, and other relatively well-resourced languages as low-resource. There is also a concrete fidelity error in Table 2. The paper provides no machine-checked proofs, reproducible code, or parameter-free derivations; its contribution is a narrative synthesis of the secondary literature.
major comments (3)
- [Section 3.1 (3.1.1 and 3.1.2)] The literature search is not described in enough detail to support the claim of a comprehensive survey. The manuscript lists broad database names and a table of keyword categories, but gives no exact query strings, no search dates, no number of records retrieved per database, no deduplication counts, and no screening counts. The exclusion criterion in Section 3.2, 'Research papers with indeterminate information were excluded,' is not defined. Because the contribution is explicitly a survey, this undocumented pipeline makes the coverage claim unfalsifiable and unreproducible. The authors should add a full search log, inclusion/exclusion criteria, and a flow diagram, or substantially soften the comprehensiveness claim to that of a selective review.
- [Section 5.2] The definition of 'low-resource' is internally inconsistent: 'except for English, we have considered all the other languages as low-resource languages' includes Spanish, French, German, Portuguese, Arabic, and Korean, several of which are among the most resourced non-English languages. This contradicts the paper's own language-distribution charts (Fig. 4), where Spanish, German, and Arabic are shown as the second-most-used languages on major platforms. The broadened scope dilutes the stated focus and changes what the survey's conclusions about 'low-resource languages' actually mean. The authors should adopt an operational definition (e.g., based on dataset size, NLP tool coverage, or speaker population) and apply it consistently throughout.
- [Table 2, row 9] The example 'People from <insert religion> should not be allowed inside.' is described in the text as religious hate speech that is neither abusive nor cyberbullying, but the table marks Religion as 'No' and Ableism as 'Yes'. This is a direct contradiction between the table and the text. The row must be corrected, and the remaining rows should be re-checked for similar inconsistencies.
minor comments (5)
- [Abstract] The abstract ends with 'Keywords: article, template, simple', which appears to be a leftover template string and should be removed before submission.
- [Section 4.2.1, Tulkens dataset] The label distribution sums incorrectly: 344 invalid comments out of 6375 is approximately 5.4%, not 0.05% as printed. Please correct the percentage and verify the arithmetic in the surrounding dataset descriptions.
- [Section 3.2] The phrase 'Research papers with indeterminate information were excluded' is vague; please specify what information was considered indeterminate and how this determination was made.
- [Section 5.2.3, Amharic] The sentence 'The Word2Vec model with Naïve Bayes classifier achieved the best performance with accuracy and ROC Score of 79.83% and 0.8305, respectively' is clear, but the preceding sentence about the Apache Spark-based model would benefit from a citation to the specific model description or the paper being summarized.
- [Section 5.2.4, Japanese subsection] The description of Kim et al. [82] refers to 'racist hate speech' measured with 'adoption thresholds of racist hate speech,' but the paper's focus on Korean racism may be clearer if the target group were named explicitly; please check that the summary matches the cited paper.
Circularity Check
No circularity: the survey synthesizes external literature and contains no derivation, fitted parameter, or prediction that reduces to its own inputs.
full rationale
This is a literature survey with no formal derivation chain, no fitted parameters, and no predictions generated from a model. Its contribution is a synthesis of externally published datasets, features, and methods, summarized in Tables 4-6 and Sections 4-5, so there is no self-referential fitting loop. The collection procedure in Section 3.1 is a keyword-based search of external databases; even if the search is underspecified or non-reproducible, that is a completeness and verifiability limitation, not circular reasoning. The scope statement in Section 5.2, 'except for English, we have considered all the other languages as low-resource languages,' is a definitional choice rather than a hidden input that forces the survey's conclusions. The Table 2 inconsistency, where a religious hate speech example is marked 'Religion: No' and 'Ableism: Yes,' is an internal data-quality or annotation error, not a reduction of a claimed result to its own assumptions. No load-bearing self-citation was identified: references to 'Das et al.' point to Mithun Das, Amit Kumar Das, and other researchers, not to the present authors, and none of the survey's organizing claims depends on a uniqueness theorem or ansatz imported from the same author group. The paper is therefore not circular.
Assumptions & free parameters
assumptions (3)
- domain assumption The hate speech definitions quoted from Meta, YouTube, X, TikTok, and LinkedIn are representative and accurately transcribed.
- domain assumption Public statistics on language usage and hate speech actions (Figs 2, 4, 6, 7) from Statista, Wikipedia, and IPSOS are accurate as cited.
- domain assumption The keyword-based search strategy in Section 3.1 yields a representative sample of the hate speech detection literature in low-resource languages.
Cite this review
Pith. "Pith review of A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages." pith.science (2026). https://pith.science/paper/N2OXGVSJ
@misc{pith2026241119017,
author = {Pith},
title = {Pith review of: A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages},
year = {2026},
howpublished = {\url{https://pith.science/paper/N2OXGVSJ}},
note = {Machine review of arXiv:2411.19017}
}
read the original abstract
The expanding influence of social media platforms over the past decade has impacted the way people communicate. The level of obscurity provided by social media and easy accessibility of the internet has facilitated the spread of hate speech. The terms and expressions related to hate speech gets updated with changing times which poses an obstacle to policy-makers and researchers in case of hate speech identification. With growing number of individuals using their native languages to communicate with each other, hate speech in these low-resource languages are also growing. Although, there is awareness about the English-related approaches, much attention have not been provided to these low-resource languages due to lack of datasets and online available data. This article provides a detailed survey of hate speech detection in low-resource languages around the world with details of available datasets, features utilized and techniques used. This survey further discusses the prevailing surveys, overlapping concepts related to hate speech, research challenges and opportunities.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Problematic content in spanish language comments in youtube videos about venezuelan refugees and migrants.Journal of Quantitative Description: Digital Media , 1, 2021
Luis Aguirre and Emese Domahidi. Problematic content in spanish language comments in youtube videos about venezuelan refugees and migrants.Journal of Quantitative Description: Digital Media , 1, 2021
2021
-
[2]
Detecting hate speech against women in english tweets
Resham Ahluwalia, Himani Soni, Edward Callow, Anderson Nascimento, and Martine De Cock. Detecting hate speech against women in english tweets. EV ALITA Evaluation of NLP and Speech Tools for Italian, 12:194, 2018
2018
-
[3]
Automatic detection of offensive language for urdu and roman urdu
Muhammad Pervez Akhter, Zheng Jiangbin, Irfan Raza Naqvi, Mohammed Abdelmajeed, and Muhammad Tariq Sadiq. Automatic detection of offensive language for urdu and roman urdu. IEEE Access, 8:91213–91226, 2020
2020
-
[4]
Detection of hate speech in social networks: a survey on multilingual corpus
Areej Al-Hassan and Hmood Al-Dossari. Detection of hate speech in social networks: a survey on multilingual corpus. In 6th international conference on computer science and information technol- ogy, volume 10, pages 10–5121. ACM, 2019
2019
-
[5]
Dataset construction for the detection of anti-social behaviour in online communication in arabic
Azalden Alakrot, Liam Murray, and Nikola S Nikolov. Dataset construction for the detection of anti-social behaviour in online communication in arabic. Procedia Computer Science, 142:174–181, 2018
2018
-
[6]
Are they our brothers? analysis and detection of religious hate speech in the arabic twittersphere
Nuha Albadi, Maram Kurdi, and Shivakant Mishra. Are they our brothers? analysis and detection of religious hate speech in the arabic twittersphere. In 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM) , pages 69–76. IEEE, 2018
2018
-
[7]
Edwin Aldana-Bobadilla, Alejandro Molina-Villegas, Yuridia Montelongo-Padilla, Ivan Lopez- Arevalo, and Oscar S. Sordia. A language model for misogyny detection in latin american spanish driven by multisource feature extraction and transformers. Applied Sciences, 11(21):10467, 2021
2021
-
[8]
Hate speech detection in the indonesian language: A dataset and preliminary study
Ika Alfina, Rio Mulia, Mohamad Ivan Fanany, and Yudo Ekanata. Hate speech detection in the indonesian language: A dataset and preliminary study. In 2017 international conference on advanced computer science and information systems (ICACSIS) , pages 233–238. IEEE, 2017
2017
Show all 160 references
-
[9]
Improving hate speech detection of urdu tweets using sentiment analysis
Muhammad Z Ali, Sahar Rauf, Kashif Javed, Sarmad Hussain, et al. Improving hate speech detection of urdu tweets using sentiment analysis. IEEE Access, 9:84296–84305, 2021
2021
-
[10]
A literature review of textual hate speech detection methods and datasets
Fatimah Alkomah and Xiaogang Ma. A literature review of textual hate speech detection methods and datasets. Information, 13(6):273, 2022
2022
-
[11]
Deep learning-based hate speech detection in code-mixed tamil text
S Anbukkarasi and S Varadhaganapathy. Deep learning-based hate speech detection in code-mixed tamil text. IETE Journal of Research , 69(11):7893–7898, 2023
2023
-
[12]
Multi- modal hate speech detection in memes using contrastive language-image pre-training.IEEE Access, 2024
Greeshma Arya, Mohammad Kamrul Hasan, Ashish Bagwari, Nurhizam Safie, Shayla Islam, Fa- tima Rayan Awad Ahmed, Aaishani De, Muhammad Attique Khan, and Taher M Ghazal. Multi- modal hate speech detection in memes using contrastive language-image pre-training.IEEE Access, 2024
2024
-
[13]
Model-agnostic meta-learning for multilingual hate speech detection
Md Rabiul Awal, Roy Ka-Wei Lee, Eshaan Tanwar, Tanmay Garg, and Tanmoy Chakraborty. Model-agnostic meta-learning for multilingual hate speech detection. IEEE Transactions on Com- putational Social Systems , 11(1):1086–1095, 2023
2023
-
[14]
Toxicity detection on bengali social media comments using supervised models
Nayan Banik and Md Hasan Hafizur Rahman. Toxicity detection on bengali social media comments using supervised models. In 2019 2nd international conference on Innovation in Engineering and Technology (ICIET), pages 1–5. IEEE, 2019
2019
-
[15]
Qutnocturnal@ hasoc’19: Cnn for hate speech and offensive content identification in hindi language
Md Abul Bashar and Richi Nayak. Qutnocturnal@ hasoc’19: Cnn for hate speech and offensive content identification in hindi language. arXiv preprint arXiv:2008.12448 , 2020
2008 arXiv
-
[16]
Semeval-2019 task 5: Multilin- gual detection of hate speech against immigrants and women in twitter
Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. Semeval-2019 task 5: Multilin- gual detection of hate speech against immigrants and women in twitter. In Proceedings of the 13t...
2019
-
[17]
Building a formal model for hate detection in french corpora
Delphine Battistelli, Cyril Bruneau, and Valentina Dragos. Building a formal model for hate detection in french corpora. Procedia Computer Science, 176:2358–2365, 2020
2020
-
[18]
Interpretable multi labeled bengali toxic comments classification using deep learning
Tanveer Ahmed Belal, GM Shahariar, and Md Hasanul Kabir. Interpretable multi labeled bengali toxic comments classification using deep learning. In 2023 International Conference on Electrical, Computer and Communication Engineering (ECCE) , pages 1–6. IEEE, 2023
2023
-
[19]
A turkish hate speech dataset and detection system
Fatih Beyhan, Buse C ¸ arık, Inan¸ c Arın, Ay¸ secan Terzio˘ glu, Berrin Yanikoglu, and Reyyan Yeniterzi. A turkish hate speech dataset and detection system. In Proceedings of the thirteenth language resources and evaluation conference, pages 4177–4185, 2022
2022
-
[20]
One to rule them all: Towards joint indic language hate speech detection
Mehar Bhatia, Tenzin Singhay Bhotia, Akshat Agarwal, Prakash Ramesh, Shubham Gupta, Kumar Shridhar, Felix Laumann, and Ayushman Dash. One to rule them all: Towards joint indic language hate speech detection. arXiv preprint arXiv:2109.13711 , 2021
2021 arXiv
-
[21]
A dataset of hindi-english code-mixed social media text for hate speech detection
Aditya Bohra, Deepanshu Vijay, Vinay Singh, Syed Sarfaraz Akhtar, and Manish Shrivastava. A dataset of hindi-english code-mixed social media text for hate speech detection. In Proceedings of the second workshop on computational modeling of people’s opinions, personality, and e...
2018
-
[22]
Hatebert: Retraining bert for abusive language detection in english
Tommaso Caselli, Valerio Basile, Jelena Mitrovi´ c, and Michael Granitzer. Hatebert: Retraining bert for abusive language detection in english. arXiv preprint arXiv:2010.12472 , 2020
2010 arXiv
-
[23]
Analyzing zero-shot transfer scenarios across spanish variants for hate speech detection
Galo Castillo-L´ opez, Arij Riabi, and Djam´ e Seddah. Analyzing zero-shot transfer scenarios across spanish variants for hate speech detection. In Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023) , pages 1–13, 2023
2023
-
[24]
Tiktok: languages covered by content moderators, 2024
Laura Ceci. Tiktok: languages covered by content moderators, 2024. URL https://www. statista.com/statistics/1405643/tiktok-language-covered-moderators/ . Accessed: 2024- 08-28
2024
-
[25]
Policycorpus xl: An italian corpus for the detection of hate speech against politics
Fabio Celli, Mirko Lai, Armend Duzha, Cristina Bosco, Viviana Patti, et al. Policycorpus xl: An italian corpus for the detection of hate speech against politics. In CEUR workshop proceedings, volume 3033, pages 1–7. CEUR-WS. org, 2021
2021
-
[26]
Offensive language identification in dravidian languages using mpnet and cnn
Bharathi Raja Chakravarthi, Manoj Balaji Jagadeeshan, Vasanth Palanikumar, and Ruba Priyad- harshini. Offensive language identification in dravidian languages using mpnet and cnn. Interna- tional Journal of Information Management Data Insights , 3(1):100151, 2023
2023
-
[27]
A twitter bert approach for offensive language detection in marathi
Tanmay Chavan, Shantanu Patankar, Aditya Kane, Omkar Gokhale, and Raviraj Joshi. A twitter bert approach for offensive language detection in marathi. arXiv preprint arXiv:2212.10039 , 2022
2022 arXiv
-
[28]
A literature survey on multimodal and multi- lingual automatic hate speech identification
Anusha Chhabra and Dinesh Kumar Vishwakarma. A literature survey on multimodal and multi- lingual automatic hate speech identification. Multimedia Systems, 29(3):1203–1230, 2023
2023
-
[29]
Conan–counter narratives through nichesourcing: a multilingual dataset of responses to fight online hate speech
Yi-Ling Chung, Elizaveta Kuzmenko, Serra Sinem Tekiroglu, and Marco Guerini. Conan–counter narratives through nichesourcing: a multilingual dataset of responses to fight online hate speech. arXiv preprint arXiv:1910.03270 , 2019
1910 arXiv
-
[30]
A corpus of turkish offensive language on social media
C ¸ a˘ grı C ¸ ¨ oltekin. A corpus of turkish offensive language on social media. InProceedings of the twelfth language resources and evaluation conference , pages 6174–6184, 2020
2020
-
[31]
Bangla hate speech detection on social media using attention-based recurrent neural network
Amit Kumar Das, Abdullah Al Asif, Anik Paul, and Md Nur Hossain. Bangla hate speech detection on social media using attention-based recurrent neural network. Journal of Intelligent Systems , 30 (1):578–591, 2021
2021
-
[32]
Data bootstrapping approaches to improve low resource abusive language detection for indic languages
Mithun Das, Somnath Banerjee, and Animesh Mukherjee. Data bootstrapping approaches to improve low resource abusive language detection for indic languages. In Proceedings of the 33rd ACM conference on hypertext and social media , pages 32–42, 2022
2022
-
[33]
Hate speech and offen- sive language detection in bengali
Mithun Das, Somnath Banerjee, Punyajoy Saha, and Animesh Mukherjee. Hate speech and offen- sive language detection in bengali. arXiv preprint arXiv:2210.03479 , 2022. 26 This work is shared under a CC BY-SA 4.0 license unless otherwise noted
2022 arXiv
-
[34]
Low-resource counterspeech generation for indic languages: The case of bengali and hindi
Mithun Das, Saurabh Kumar Pandey, Shivansh Sethi, Punyajoy Saha, and Animesh Mukherjee. Low-resource counterspeech generation for indic languages: The case of bengali and hindi. arXiv preprint arXiv:2402.07262, 2024
2024 arXiv
-
[35]
Automated hate speech de- tection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. Automated hate speech de- tection and the problem of offensive language. In Proceedings of the international AAAI conference on web and social media , volume 11, pages 512–515, 2017
2017
-
[36]
Sentiment analysis methods for politics and hate speech contents in spanish language: a systematic review
Ernesto del Valle and Luis de la Fuente. Sentiment analysis methods for politics and hate speech contents in spanish language: a systematic review. IEEE Latin America Transactions , 21(3): 408–418, 2023
2023
-
[37]
Detox: A comprehensive dataset for german offensive language and conversation analysis
Christoph Demus, Jonas Pitz, Mina Sch¨ utz, Nadine Probol, Melanie Siegel, and Dirk Labudde. Detox: A comprehensive dataset for german offensive language and conversation analysis. In Proceedings of the Sixth Workshop on Online Abuse and Harms (WOAH) , pages 143–153, 2022
2022
-
[38]
Statista dataset for facebook, 2023
Statista Research Department. Statista dataset for facebook, 2023. URL https://www.statista. com/statistics/1013804/facebook-hate-speech-content-deletion-quarter/ . Accessed: 2024-08-28
2023
-
[39]
Hate speech detection in asian languages: a survey
LK Dhanya and Kannan Balakrishnan. Hate speech detection in asian languages: a survey. In 2021 international conference on communication, control and information sciences (ICCISc) , volume 1, pages 1–5. IEEE, 2021
2021
-
[40]
Com- mon sense reasoning for detection, prevention, and mitigation of cyberbullying
Karthik Dinakar, Birago Jones, Catherine Havasi, Henry Lieberman, and Rosalind Picard. Com- mon sense reasoning for detection, prevention, and mitigation of cyberbullying. ACM Transactions on Interactive Intelligent Systems (TiiS) , 2(3):1–30, 2012
2012
-
[41]
Statista dataset for instagram, 2023
Stacy Jo Dixon. Statista dataset for instagram, 2023. URL https://www.statista.com/ statistics/1275933/global-actioned-hate-speech-content-instagram/ . Accessed: 2024- 08-28
2023
-
[42]
Statista most used social networks, 2024
Stacy Jo Dixon. Statista most used social networks, 2024. URL https://www.statista. com/statistics/272014/global-social-networks-ranked-by-number-of-users/ . Accessed: 2024-08-28
2024
-
[43]
Hate speech detection on vietnamese social media text using the bidirectional-lstm model
Hang Thi-Thuy Do, Huy Duc Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen, and Anh Gia- Tuan Nguyen. Hate speech detection on vietnamese social media text using the bidirectional-lstm model. arXiv preprint arXiv:1911.03648 , 2019
1911 arXiv
-
[44]
At the lower end of language—exploring the vulgar and obscene side of german
Elisabeth Eder, Ulrike Krieg-Holz, and Udo Hahn. At the lower end of language—exploring the vulgar and obscene side of german. In Proceedings of the third workshop on abusive language online, pages 119–128, 2019
2019
-
[45]
A deep learning approach to detect abusive bengali text
Estiak Ahmed Emon, Shihab Rahman, Joti Banarjee, Amit Kumar Das, and Tanni Mittra. A deep learning approach to detect abusive bengali text. In 2019 7th International Conference on Smart Computing & Communications (ICSCC) , pages 1–5. IEEE, 2019
2019
-
[46]
A study on the feasibility to detect hate speech in swedish
Johan Fernquist, Oskar Lindholm, Lisa Kaati, and Nazar Akrami. A study on the feasibility to detect hate speech in swedish. In 2019 IEEE international conference on big data (Big Data) , pages 4724–4729. IEEE, 2019
2019
-
[47]
A survey on automatic detection of hate speech in text
Paula Fortuna and S´ ergio Nunes. A survey on automatic detection of hate speech in text. ACM Computing Surveys (CSUR) , 51(4):1–30, 2018
2018
-
[48]
A hierarchically-labeled portuguese hate speech dataset
Paula Fortuna, Joao Rocha da Silva, Leo Wanner, S´ ergio Nunes, et al. A hierarchically-labeled portuguese hate speech dataset. In Proceedings of the third workshop on abusive language online , pages 94–104, 2019
2019
-
[49]
Defining disability: Understandings of and attitudes towards ableism and disability
Carli Friedman and Aleksa L Owen. Defining disability: Understandings of and attitudes towards ableism and disability. Disability Studies Quarterly , 37(1), 2017
2017
-
[50]
Normalizing misogyny: hate speech and verbal abuse of female politicians on japanese twitter
Tamara Fuchs and Fabian Sch¨ afer. Normalizing misogyny: hate speech and verbal abuse of female politicians on japanese twitter. In Japan forum, volume 33, pages 553–579. Taylor & Francis, 2021. 27 This work is shared under a CC BY-SA 4.0 license unless otherwise noted
2021
-
[51]
Cross- lingual offensive language identification for low resource languages: The case of marathi
Saurabh Gaikwad, Tharindu Ranasinghe, Marcos Zampieri, and Christopher M Homan. Cross- lingual offensive language identification for low resource languages: The case of marathi. arXiv preprint arXiv:2109.03552, 2021
2021 arXiv
-
[52]
Hate speech detection: a comparison of mono and multi- lingual transformer model with cross-language evaluation
Koyel Ghosh and Apurbalal Senapati. Hate speech detection: a comparison of mono and multi- lingual transformer model with cross-language evaluation. In Proceedings of the 36th Pacific Asia Conference on Language, Information and Computation , pages 853–865, 2022
2022
-
[53]
Annihilate hates (task 4 hasoc 2023): Hate speech detection in assamese bengali and bodo languages
Koyel Ghosh, Apurbalal Senapati, and Aditya Shankar Pal. Annihilate hates (task 4 hasoc 2023): Hate speech detection in assamese bengali and bodo languages. In FIRE (Working Notes) , pages 368–382, 2023
2023
-
[54]
Transformer-based hate speech detection in assamese
Koyel Ghosh, Debarshi Sonowal, Abhilash Basumatary, Bidisha Gogoi, and Apurbalal Senapati. Transformer-based hate speech detection in assamese. In 2023 IEEE Guwahati Subsection Confer- ence (GCON), pages 1–5. IEEE, 2023
2023
-
[55]
Sehc: A benchmark setup to identify online hate speech in english
Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya, Tista Saha, Alka Kumar, and Shikha Sri- vastava. Sehc: A benchmark setup to identify online hate speech in english. IEEE Transactions on Computational Social Systems , 10(2):760–770, 2022
2022
-
[56]
Exploring hate speech detection in multimodal publications
Raul Gomez, Jaume Gibert, Lluis Gomez, and Dimosthenis Karatzas. Exploring hate speech detection in multimodal publications. In Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages 1470–1478, 2020
2020
-
[57]
Multilingual abusive comment detection at scale for indic languages
Vikram Gupta, Sumegh Roychowdhury, Mithun Das, Somnath Banerjee, Punyajoy Saha, Binny Mathew, Animesh Mukherjee, et al. Multilingual abusive comment detection at scale for indic languages. Advances in Neural Information Processing Systems , 35:26176–26191, 2022
2022
-
[58]
Kancmd: Kannada codemixed dataset for sentiment analysis and offensive language detection
Adeep Hande, Ruba Priyadharshini, and Bharathi Raja Chakravarthi. Kancmd: Kannada codemixed dataset for sentiment analysis and offensive language detection. In Proceedings of the Third Workshop on Computational Modeling of People’s Opinions, Personality, and Emotion ’s in Soci...
2020
-
[59]
Multi-class sentiment classification on bengali social media comments using machine learning
Rezaul Haque, Naimul Islam, Mayisha Tasneem, and Amit Kumar Das. Multi-class sentiment classification on bengali social media comments using machine learning. International journal of cognitive computing in engineering , 4:21–35, 2023
2023
-
[60]
Multi-class hate speech detection in the norwegian language using fast-rnn and multilingual fine-tuned transformers
Ehtesham Hashmi and Sule Yildirim Yayilgan. Multi-class hate speech detection in the norwegian language using fast-rnn and multilingual fine-tuned transformers. Complex & Intelligent Systems , 10(3):4535–4556, 2024
2024
-
[61]
Mute: A multimodal dataset for detecting hateful memes
Eftekhar Hossain, Omar Sharif, and Mohammed Moshiul Hoque. Mute: A multimodal dataset for detecting hateful memes. In Proceedings of the 2nd conference of the asia-pacific chapter of the association for computational linguistics and the 12th international joint conference on n...
2022
-
[62]
Baad: A multipurpose dataset for automatic bangla offensive speech recognition
Md Fahad Hossain, Md Al Abid Supto, Zannat Chowdhury, Hana Sultan Chowdhury, and Sheikh Abujar. Baad: A multipurpose dataset for automatic bangla offensive speech recognition. Data in Brief, 48:109067, 2023
2023
-
[63]
Survey on the impact of online disinformation and hate speech, 2023
IPSOS. Survey on the impact of online disinformation and hate speech, 2023. URL https://www.ipsos.com/sites/default/files/ct/news/documents/2023-11/ unesco-ipsos-online-disinformation-hate-speech.pdf . Accessed: 2024-08-28
2023
-
[64]
Hateful speech detection in public facebook pages for the bengali language
Alvi Md Ishmam and Sadia Sharmin. Hateful speech detection in public facebook pages for the bengali language. In 2019 18th IEEE international conference on machine learning and applications (ICMLA), pages 555–560. IEEE, 2019
2019
-
[65]
A systematic review of hate speech automatic detection using natural language processing
Md Saroar Jahan and Mourad Oussalah. A systematic review of hate speech automatic detection using natural language processing. Neurocomputing, page 126232, 2023
2023
-
[66]
Banglahatebert: Bert for abusive language detection in bengali
Md Saroar Jahan, Mainul Haque, Nabil Arhab, and Mourad Oussalah. Banglahatebert: Bert for abusive language detection in bengali. In Proceedings of the second international workshop on resources and techniques for user information in abusive language analysis , pages 8–15, 2022...
2022
-
[67]
Finnish hate-speech detection on social media using cnn and finbert
Md Saroar Jahan, Mourad Oussalah, and Nabil Arhab. Finnish hate-speech detection on social media using cnn and finbert. In Language Resources and Evaluation Conference, LREC 2022, 20-25 June 2022, Palais du Pharo, Marseille, France: conference proceedings . European Language R...
2022
-
[68]
Right-wing german hate speech on twitter: Analysis and automatic detection
Sylvia Jaki and Tom De Smedt. Right-wing german hate speech on twitter: Analysis and automatic detection. arXiv preprint arXiv:1910.07518 , 2019
1910 arXiv
-
[69]
Toxicity detection for indic multilingual social media content
Manan Jhaveri, Devanshu Ramaiya, and Harveen Singh Chadha. Toxicity detection for indic multilingual social media content. arXiv preprint arXiv:2201.00598 , 2022
2022 arXiv
-
[70]
Harnessing pre-trained sentence transformers for offensive lan- guage detection in indian languages
Ananya Joshi and Raviraj Joshi. Harnessing pre-trained sentence transformers for offensive lan- guage detection in indian languages. arXiv preprint arXiv:2310.02249 , 2023
2023 arXiv
-
[71]
Bangla hate speech de- tection in videos using machine learning
Mohd Istiaq Hossain Junaid, Faisal Hossain, and Rashedur M Rahman. Bangla hate speech de- tection in videos using machine learning. In 2021 IEEE 12th Annual Ubiquitous Computing, Elec- tronics & Mobile Communication Conference (UEMCON) , pages 0347–0351. IEEE, 2021
2021
-
[72]
Korean online hate speech dataset for multilabel classification: How can social science improve dataset on hate speech? arXiv preprint arXiv:2204.03262 , 2022
TaeYoung Kang, Eunrang Kwon, Junbum Lee, Youngeun Nam, Junmo Song, and JeongKyu Suh. Korean online hate speech dataset for multilabel classification: How can social science improve dataset on hate speech? arXiv preprint arXiv:2204.03262 , 2022
2022 arXiv
-
[73]
Hatespeech and offensive content detection in hindi language using c-bigru
Sudharsana Kannan and Jelena Mitrovic. Hatespeech and offensive content detection in hindi language using c-bigru. In FIRE (Working Notes) , pages 209–216, 2021
2021
-
[74]
Classifica- tion benchmarks for under-resourced bengali language based on multichannel convolutional-lstm network
Md Rezaul Karim, Bharathi Raja Chakravarthi, John P McCrae, and Michael Cochez. Classifica- tion benchmarks for under-resourced bengali language based on multichannel convolutional-lstm network. In 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (...
2020
-
[75]
Deephateexplainer: Explainable hate speech detection in under-resourced bengali language
Md Rezaul Karim, Sumon Kanti Dey, Tanhim Islam, Sagor Sarker, Mehadi Hasan Menon, Kabir Hossain, Md Azam Hossain, and Stefan Decker. Deephateexplainer: Explainable hate speech detection in under-resourced bengali language. In 2021 IEEE 8th international conference on data scie...
2021
-
[76]
Multimodal hate speech detection from bengali memes and texts
Md Rezaul Karim, Sumon Kanti Dey, Tanhim Islam, Md Shajalal, and Bharathi Raja Chakravarthi. Multimodal hate speech detection from bengali memes and texts. In International Conference on Speech and Language Technologies for Low-resource Languages , pages 293–308. Springer, 2022
2022
-
[77]
G-bert: an efficient method for identifying hate speech in bengali texts on social media
Ashfia Jannat Keya, Md Mohsin Kabir, Nusrat Jahan Shammey, MF Mridha, Md Rashedul Islam, and Yutaka Watanobe. G-bert: an efficient method for identifying hate speech in bengali texts on social media. IEEE Access, 2023
2023
-
[78]
Offensive language detection for low resource language using deep sequence model
Anas Ali Khan, M Hammad Iqbal, Shibli Nisar, Awais Ahmad, and Waseem Iqbal. Offensive language detection for low resource language using deep sequence model. IEEE Transactions on Computational Social Systems , 2023
2023
-
[79]
Muril: Multilin- gual representations for indian languages
Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, et al. Muril: Multilin- gual representations for indian languages. arXiv preprint arXiv:2103.10730 , 2021
2021 arXiv
-
[80]
Animojity: Detecting hate comments in indic languages and analysing bias against content creators
Rahul Khurana, Chaitanya Pandey, Priyanshi Gupta, and Preeti Nagrath. Animojity: Detecting hate comments in indic languages and analysing bias against content creators. In Proceedings of the 19th International Conference on Natural Language Processing (ICON) , pages 172–182, 2022
2022
-
[81]
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ring- shia, and Davide Testuggine. The hateful memes challenge: Detecting hate speech in multimodal memes. Advances in neural information processing systems , 33:2611–2624, 2020
2020
-
[82]
The impact of politicians’ behaviors on hate speech spread: hate speech adoption threshold on twitter in japan
Taehee Kim and Yuki Ogawa. The impact of politicians’ behaviors on hate speech spread: hate speech adoption threshold on twitter in japan. Journal of Computational Social Science , pages 1–26, 2024. 29 This work is shared under a CC BY-SA 4.0 license unless otherwise noted
2024
-
[83]
Exploiting unsupervised pre-training and automated feature engineering for low-resource hate speech detection in polish
Renard Korzeniowski, Rafa l Rolczy´ nski, Przemys law Sadownik, Tomasz Korbak, and Marcin Mo˙ zejko. Exploiting unsupervised pre-training and automated feature engineering for low-resource hate speech detection in polish. arXiv preprint arXiv:1906.09325 , 2019
1906 arXiv
-
[84]
Hate speech detection in algerian dialect using deep learning
Dihia Lanasri, Juan Olano, Sifal Klioui, Sin Liang Lee, and Lamia Sekkai. Hate speech detection in algerian dialect using deep learning. arXiv preprint arXiv:2309.11611 , 2023
2023 arXiv
-
[85]
Project hatemeter: helping ngos and social science researchers to analyze and prevent anti-muslim hate speech on social media
Mario Laurent. Project hatemeter: helping ngos and social science researchers to analyze and prevent anti-muslim hate speech on social media. Procedia Computer Science , 176:2143–2153, 2020
2020
-
[86]
K-mhas: A multi-label hate speech detection dataset in korean online news comment
Jean Lee, Taejun Lim, Heejun Lee, Bogeun Jo, Yangsok Kim, Heegeun Yoon, and Soyeon Caren Han. K-mhas: A multi-label hate speech detection dataset in korean online news comment. arXiv preprint arXiv:2208.10684, 2022
2022 arXiv
-
[87]
Crehate: Cross-cultural re-annotation of english hate speech dataset
Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, Juho Kim, and Alice Oh. Crehate: Cross-cultural re-annotation of english hate speech dataset. arXiv preprint arXiv:2308.16705 , 2023
2023 arXiv
-
[88]
Toxic language detection in social media for brazilian portuguese: New dataset and multilingual analysis
Joao A Leite, Diego F Silva, Kalina Bontcheva, and Carolina Scarton. Toxic language detection in social media for brazilian portuguese: New dataset and multilingual analysis. arXiv preprint arXiv:2010.04543, 2020
2010 arXiv
-
[89]
A large-scale dataset for hate speech detection on vietnamese social media texts
Son T Luu, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. A large-scale dataset for hate speech detection on vietnamese social media texts. In Advances and Trends in Artificial Intelligence. Artificial Intelligence Practices: 34th International Conference on Industrial, Engineerin...
2021
-
[90]
Cyberbullying detection for low-resource languages and dialects: Review of the state of the art
Tanjim Mahmud, Michal Ptaszynski, Juuso Eronen, and Fumito Masui. Cyberbullying detection for low-resource languages and dialects: Review of the state of the art. Information Processing & Management, 60(5):103454, 2023
2023
-
[91]
Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in indo-european languages
Thomas Mandl, Sandip Modha, Prasenjit Majumder, Daksh Patel, Mohana Dave, Chintak Man- dlia, and Aditya Patel. Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in indo-european languages. In Proceedings of the 11th annual meeting of th...
2019
-
[92]
Tueval at semeval-2019 task 5: Lstm approach to hate speech detection in english and spanish
Mihai Manolescu, Denise L¨ offlad, Adham Nasser Mohamed Saber, and Masoumeh Moradipour Tari. Tueval at semeval-2019 task 5: Lstm approach to hate speech detection in english and spanish. In Proceedings of the 13th International Workshop on Semantic Evaluation , pages 498– 502, 2019
2019
-
[93]
Twitter hate speech detection: a systematic review of methods, taxonomy analysis, challenges, and opportunities
Zainab Mansur, Nazlia Omar, and Sabrina Tiun. Twitter hate speech detection: a systematic review of methods, taxonomy analysis, challenges, and opportunities. IEEE Access, 11:16226– 16249, 2023
2023
-
[94]
An ensemble approach for dutch cross-domain hate speech detection
Ilia Markov, Ine Gevers, and Walter Daelemans. An ensemble approach for dutch cross-domain hate speech detection. In International conference on applications of natural language to information systems, pages 3–15. Springer, 2022
2022
-
[95]
Spread of hate speech in online social media
Binny Mathew, Ritam Dutt, Pawan Goyal, and Animesh Mukherjee. Spread of hate speech in online social media. In Proceedings of the 10th ACM conference on web science , pages 173–182, 2019
2019
-
[96]
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. Hatexplain: A benchmark dataset for explainable hate speech detection. In Proceedings of the AAAI conference on artificial intelligence , volume 35, pages 14867–14875, 2021
2021
-
[97]
Did you offend me? classification of offensive tweets in hinglish language
Puneet Mathur, Ramit Sawhney, Meghna Ayyar, and Rajiv Shah. Did you offend me? classification of offensive tweets in hinglish language. In Proceedings of the 2nd workshop on abusive language online (AL W2), pages 138–148, 2018. 30 This work is shared under a CC BY-SA 4.0 licen...
2018
-
[98]
Overview of the hasoc subtrack at fire 2021: Hate speech and offensive content identification in english and indo-aryan languages and conversational hate speech
Sandip Modha, Thomas Mandl, Gautam Kishore Shahi, Hiren Madhu, Shrey Satapara, Tharindu Ranasinghe, and Marcos Zampieri. Overview of the hasoc subtrack at fire 2021: Hate speech and offensive content identification in english and indo-aryan languages and conversational hate sp...
2021
-
[99]
Ethos: a multi- label hate speech detection dataset
Ioannis Mollas, Zoe Chrysopoulou, Stamatis Karlos, and Grigorios Tsoumakas. Ethos: a multi- label hate speech detection dataset. Complex & Intelligent Systems , 8(6):4663–4678, 2022
2022
-
[100]
Social network hate speech detection for amharic language
Zewdie Mossie, Jenq-Haur Wang, et al. Social network hate speech detection for amharic language. Computer Science & Information Technology , pages 41–55, 2018
2018
-
[101]
Abusive language detection on arabic social media
Hamdy Mubarak, Kareem Darwish, and Walid Magdy. Abusive language detection on arabic social media. In Proceedings of the first workshop on abusive language online , pages 52–56, 2017
2017
-
[102]
Hate speech and offensive content detection in indo-aryan languages: A battle of lstm and transformers
Nikhil Narayan, Mrutyunjay Biswal, Pramod Goyal, and Abhranta Panigrahi. Hate speech and offensive content detection in indo-aryan languages: A battle of lstm and transformers. arXiv preprint arXiv:2312.05671, 2023
2023 arXiv
-
[103]
Offensive language detection in nepali social media
Nobal B Niraula, Saurab Dulal, and Diwa Koirala. Offensive language detection in nepali social media. In Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021) , pages 67–75, 2021
2021
-
[104]
Tackling hate speech in low-resource languages with context experts
Daniel Nkemelu, Harshil Shah, Michael Best, and Irfan Essa. Tackling hate speech in low-resource languages with context experts. InProceedings of the 2022 International Conference on Information and Communication Technologies and Development , pages 1–11, 2022
2022
-
[105]
Abusive language detection in online user content
Chikashi Nobata, Joel Tetreault, Achint Thomas, Yashar Mehdad, and Yi Chang. Abusive language detection in online user content. In Proceedings of the 25th international conference on world wide web, pages 145–153, 2016
2016
-
[106]
Encyclopedia of gender and society , volume 2
Jodi O’Brien. Encyclopedia of gender and society , volume 2. Sage, 2009
2009
-
[107]
Evaluating machine learning techniques for detecting offensive and hate speech in south african tweets
Oluwafemi Oriola and Eduan Kotz´ e. Evaluating machine learning techniques for detecting offensive and hate speech in south african tweets. IEEE Access, 8:21496–21509, 2020
2020
-
[108]
Multilin- gual and multi-aspect hate speech analysis
Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang, Yangqiu Song, and Dit-Yan Yeung. Multilin- gual and multi-aspect hate speech analysis. arXiv preprint arXiv:1908.11049 , 2019
1908 arXiv
-
[109]
Austrotox: A dataset for target-based austrian german offensive language detection
Pia Pachinger, Janis Goldzycher, Anna Maria Planitzer, Wojciech Kusa, Allan Hanbury, and Julia Neidhardt. Austrotox: A dataset for target-based austrian german offensive language detection. arXiv preprint arXiv:2406.08080 , 2024
2024 arXiv
-
[110]
Towards multidomain and multi- lingual abusive language detection: a survey
Endang Wahyu Pamungkas, Valerio Basile, and Viviana Patti. Towards multidomain and multi- lingual abusive language detection: a survey. Personal and Ubiquitous Computing , 27(1):17–43, 2023
2023
-
[111]
Regional language toxic comment classification
Yashkumar Parikh and Jinan Fiaidhi. Regional language toxic comment classification. Authorea Preprints, 2023
2023
-
[112]
Multimodal hate speech detection in greek social media
Konstantinos Perifanos and Dionysis Goutsos. Multimodal hate speech detection in greek social media. Multimodal Technologies and Interaction, 5(7):34, 2021
2021
-
[113]
Offensive language identification in greek
Zeses Pitenis, Marcos Zampieri, and Tharindu Ranasinghe. Offensive language identification in greek. arXiv preprint arXiv:2003.07459 , 2020
2003 arXiv
-
[114]
Comparing pre-trained language models for spanish hate speech detection.Expert Systems with Applications , 166:114120, 2021
Flor Miriam Plaza-del Arco, M Dolores Molina-Gonz´ alez, L Alfonso Urena-L´ opez, and M Teresa Mart ´ ın-Valdivia. Comparing pre-trained language models for spanish hate speech detection.Expert Systems with Applications , 166:114120, 2021
2021
-
[115]
Hate speech annotation: Analysis of an italian twitter corpus
Fabio Poletto, Marco Stranisci, Manuela Sanguinetti, Viviana Patti, Cristina Bosco, et al. Hate speech annotation: Analysis of an italian twitter corpus. In Ceur workshop proceedings, volume 2006, pages 1–6. CEUR-WS, 2017. 31 This work is shared under a CC BY-SA 4.0 license un...
2006
-
[116]
Results of the poleval 2019 shared task 6: First dataset and open shared task for automatic cyberbullying detection in polish twitter
Michal Ptaszynski, Agata Pieciukiewicz, and Pawe l Dyba la. Results of the poleval 2019 shared task 6: First dataset and open shared task for automatic cyberbullying detection in polish twitter. 2019
2019
-
[117]
A benchmark dataset for learning to intervene in online hate speech
Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding, and William Yang Wang. A benchmark dataset for learning to intervene in online hate speech. arXiv preprint arXiv:1909.04251 , 2019
1909 arXiv
-
[118]
Sold: Sinhala offensive language dataset
Tharindu Ranasinghe, Isuri Anuradha, Damith Premasiri, Kanishka Silva, Hansi Hettiarachchi, Lasitha Uyangodage, and Marcos Zampieri. Sold: Sinhala offensive language dataset. Language Resources and Evaluation, pages 1–41, 2024
2024
-
[119]
A comparative study of different state-of-the-art hate speech detection methods in hindi-english code-mixed data
Priya Rani, Shardul Suryawanshi, Koustava Goswami, Bharathi Raja Chakravarthi, Theodorus Fransen, and John Philip McCrae. A comparative study of different state-of-the-art hate speech detection methods in hindi-english code-mixed data. In Proceedings of the second workshop on ...
2020
-
[120]
Toxlex bn: A curated dataset of bangla toxic language derived from facebook comment
Mohammad Mamun Or Rashid. Toxlex bn: A curated dataset of bangla toxic language derived from facebook comment. Data in Brief , 43:108416, 2022
2022
-
[121]
Offensive language detection using multi-level classification
Amir H Razavi, Diana Inkpen, Sasha Uritsky, and Stan Matwin. Offensive language detection using multi-level classification. In Advances in Artificial Intelligence: 23rd Canadian Conference on Artificial Intelligence, Canadian AI 2010, Ottawa, Canada, May 31–June 2, 2010. Proce...
2010
-
[122]
Hate speech detection in the bengali language: A dataset and its baseline evaluation
Nauros Romim, Mosahed Ahmed, Hriteshwar Talukder, and Md Saiful Islam. Hate speech detection in the bengali language: A dataset and its baseline evaluation. In Proceedings of International Joint Conference on Advances in Computational Intelligence: IJCACI 2020 , pages 457–468....
2020
-
[123]
Bd-shs: A benchmark dataset for learning to detect online bangla hate speech in different social contexts
Nauros Romim, Mosahed Ahmed, Md Saiful Islam, Arnab Sen Sharma, Hriteshwar Talukder, and Mohammad Ruhul Amin. Bd-shs: A benchmark dataset for learning to detect online bangla hate speech in different social contexts. arXiv preprint arXiv:2206.00372 , 2022
2022 arXiv
-
[124]
Measuring the reliability of hate speech annotations: The case of the european refugee crisis
Bj¨ orn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki. Measuring the reliability of hate speech annotations: The case of the european refugee crisis. arXiv preprint arXiv:1701.08118 , 2017
2017 arXiv
-
[125]
Hatecheck: Functional tests for hate speech detection models
Paul R¨ ottger, Bertram Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet B Pierrehumbert. Hatecheck: Functional tests for hate speech detection models. arXiv preprint arXiv:2012.15606, 2020
2012 arXiv
-
[126]
A framework for hate speech detection using deep convolutional neural network
Pradeep Kumar Roy, Asis Kumar Tripathy, Tapan Kumar Das, and Xiao-Zhi Gao. A framework for hate speech detection using deep convolutional neural network. IEEE Access, 8:204951–204962, 2020
2020
-
[127]
Hate speech and offensive language detection in dravidian languages using deep ensemble framework
Pradeep Kumar Roy, Snehaan Bhawal, and Chinnaudayar Navaneethakrishnan Subalalitha. Hate speech and offensive language detection in dravidian languages using deep ensemble framework. Computer Speech & Language , 75:101386, 2022
2022
-
[128]
Detection of hate speech using bert and hate speech word embedding with deep model
Hind Saleh, Areej Alhothali, and Kawthar Moria. Detection of hate speech using bert and hate speech word embedding with deep model. Applied Artificial Intelligence , 37(1):2166719, 2023
2023
-
[129]
Sinhala hate speech detection in social media using text mining and machine learning
HMST Sandaruwan, SAS Lorensuhewa, and MAL Kalyani. Sinhala hate speech detection in social media using text mining and machine learning. In 2019 19th International Conference on Advances in ICT for Emerging Regions (ICTer) , volume 250, pages 1–8. IEEE, 2019
2019
-
[130]
An italian twitter corpus of hate speech against immigrants
Manuela Sanguinetti, Fabio Poletto, Cristina Bosco, Viviana Patti, and Marco Stranisci. An italian twitter corpus of hate speech against immigrants. In Proceedings of the eleventh international conference on language resources and evaluation (LREC 2018) , 2018
2018
-
[131]
Facebook’s top ten languages, 2018
Juan Pablo Sans. Facebook’s top ten languages, 2018. URL https://www.linkedin.com/pulse/ facebooks-top-ten-languages-who-using-them-juan-pablo# . Accessed: 2024-08-28. 32 This work is shared under a CC BY-SA 4.0 license unless otherwise noted
2018
-
[132]
Identifying vulgarity in bengali social media textual content
Salim Sazzed. Identifying vulgarity in bengali social media textual content. PeerJ Computer Science, 7:e665, 2021
2021
-
[133]
Top languages on twitter-stats, 2024
Semiocast. Top languages on twitter-stats, 2024. URL https://semiocast.com/ top-languages-on-twitter-stats/# . Accessed: 2024-08-28
2024
-
[134]
Thar-targeted hate speech against reli- gion: A high-quality hindi-english code-mixed dataset with the application of deep learning models for automatic detection
Deepawali Sharma, Aakash Singh, and Vivek Kumar Singh. Thar-targeted hate speech against reli- gion: A high-quality hindi-english code-mixed dataset with the application of deep learning models for automatic detection. ACM Transactions on Asian and Low-Resource Language Inform...
2024
-
[135]
Targets and aspects in social media hate speech
Alexander Shvets, Paula Fortuna, Juan Soler, and Leo Wanner. Targets and aspects in social media hate speech. In Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021), pages 179–190, 2021
2021
-
[136]
Offensive language and hate speech detection for danish
Gudbjartur Ingi Sigurbergsson and Leon Derczynski. Offensive language and hate speech detection for danish. arXiv preprint arXiv:1908.04531 , 2019
1908 arXiv
-
[137]
Cyberbullying: Its nature and impact in secondary school pupils
Peter K Smith, Jess Mahdavi, Manuel Carvalho, Sonja Fisher, Shanette Russell, and Neil Tippett. Cyberbullying: Its nature and impact in secondary school pupils. Journal of child psychology and psychiatry, 49(4):376–385, 2008
2008
-
[138]
Detection of hate speech text in hindi-english code- mixed data
K Sreelakshmi, B Premjith, and KP Soman. Detection of hate speech text in hindi-english code- mixed data. Procedia Computer Science, 171:737–744, 2020
2020
-
[139]
Offensive language detection in tamil youtube comments by adapters and cross-domain knowledge transfer
Malliga Subramanian, Rahul Ponnusamy, Sean Benhur, Kogilavani Shanmugavadivel, Adhithiya Ganesan, Deepti Ravi, Gowtham Krishnan Shanmugasundaram, Ruba Priyadharshini, and Bharathi Raja Chakravarthi. Offensive language detection in tamil youtube comments by adapters and cross-d...
2022
-
[140]
Detection of abusive bengali comments for mixed social media data using machine learning
Sherin Sultana, Md Omur Faruk Redoy, Jabir Al Nahian, Abu Kaisar Mohammad Masum, and Sheikh Abujar. Detection of abusive bengali comments for mixed social media data using machine learning. 2023
2023
-
[141]
Indonesia hate speech detection using deep learning
Taufic Leonardo Sutejo and Dessi Puji Lestari. Indonesia hate speech detection using deep learning. In 2018 International Conference on Asian Language Processing (IALP), pages 39–43. IEEE, 2018
2018
-
[142]
Automated amharic hate speech posts and comments detection model using recurrent neural network
Surafel Getachew Tesfaye and Kula Kakeba. Automated amharic hate speech posts and comments detection model using recurrent neural network. 2020
2020
-
[143]
A multi-modal dataset for hate speech detection on social media: Case-study of russia-ukraine conflict
Surendrabikram Thapa, Aditya Shah, Farhan Ahmad Jafri, Usman Naseem, and Imran Razzak. A multi-modal dataset for hate speech detection on social media: Case-study of russia-ukraine conflict. In CASE 2022-5th Workshop on Challenges and Applications of Automated Extraction of So...
2022
-
[144]
A dictionary-based approach to racism detection in dutch social media
St´ ephan Tulkens, Lisa Hilte, Elise Lodewyckx, Ben Verhoeven, and Walter Daelemans. A dictionary-based approach to racism detection in dutch social media. arXiv preprint arXiv:1608.08738, 2016
2016 arXiv
-
[145]
Quarantining online hate speech: technical and ethical perspectives
Stefanie Ullmann and Marcus Tomalin. Quarantining online hate speech: technical and ethical perspectives. Ethics and Information Technology , 22(1):69–80, 2020
2020
-
[146]
Guidelines for the fine-grained analysis of cyberbullying
Cynthia Van Hee, Ben Verhoeven, Els Lefever, Guy De Pauw, V´ eronique Hoste, and Walter Daele- mans. Guidelines for the fine-grained analysis of cyberbullying. 2015
2015
-
[147]
Detection of racist language in french tweets
Natalia Vanetik and Elisheva Mimoun. Detection of racist language in french tweets. Information, 13(7):318, 2022
2022
-
[148]
Hatebr: A large expert annotated corpus of brazilian instagram comments for offensive language and hate speech detection.arXiv preprint arXiv:2103.14972, 2021
Francielle Alves Vargas, Isabelle Carvalho, Fabiana Rodrigues de G´ oes, Fabr ´ ıcio Benevenuto, and Thiago Alexandre Salgueiro Pardo. Hatebr: A large expert annotated corpus of brazilian instagram comments for offensive language and hate speech detection.arXiv preprint arXiv:...
2021 arXiv
-
[149]
Towards offensive language identifica- tion for tamil code-mixed youtube comments and posts
Charangan Vasantharajan and Uthayasanker Thayasivam. Towards offensive language identifica- tion for tamil code-mixed youtube comments and posts. SN Computer Science , 3(1):94, 2022. 33 This work is shared under a CC BY-SA 4.0 license unless otherwise noted
2022
-
[150]
Online multilingual hate speech detection: experimenting with hindi and english social media
Neeraj Vashistha and Arkaitz Zubiaga. Online multilingual hate speech detection: experimenting with hindi and english social media. Information, 12(1):5, 2020
2020
-
[151]
Hate and offensive speech detection in hindi and marathi
Abhishek Velankar, Hrushikesh Patil, Amol Gore, Shubham Salunke, and Raviraj Joshi. Hate and offensive speech detection in hindi and marathi. arXiv preprint arXiv:2110.12200 , 2021
2021 arXiv
-
[152]
L3cube- mahahate: A tweet-based marathi hate speech detection dataset and bert models
Abhishek Velankar, Hrushikesh Patil, Amol Gore, Shubham Salunke, and Raviraj Joshi. L3cube- mahahate: A tweet-based marathi hate speech detection dataset and bert models. arXiv preprint arXiv:2203.13778, 2022
2022 arXiv
-
[153]
Mono vs multilingual bert for hate speech detection and text classification: A case study in marathi
Abhishek Velankar, Hrushikesh Patil, and Raviraj Joshi. Mono vs multilingual bert for hate speech detection and text classification: A case study in marathi. In IAPR Workshop on Artificial Neural Networks in Pattern Recognition , pages 121–128. Springer, 2022
2022
-
[154]
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy. Hateful symbols or hateful people? predictive features for hate speech detection on twitter. In Proceedings of the NAACL student research workshop, pages 88–93, 2016
2016
-
[155]
German abusive language dataset with focus on covid-19
Maximilian Wich, Svenja R¨ ather, and Georg Groh. German abusive language dataset with focus on covid-19. In Proceedings of the 17th Conference on Natural Language Processing (KONVENS 2021), pages 247–252, 2021
2021
-
[156]
Overview of the germeval 2018 shared task on the identification of offensive language
Michael Wiegand, Melanie Siegel, and Josef Ruppenhofer. Overview of the germeval 2018 shared task on the identification of offensive language. 2018
2018
-
[157]
Content languages on youtube, 2024
Wikipedia. Content languages on youtube, 2024. URL https://en.wikipedia.org/wiki/ Languages_used_on_the_Internet. Accessed: 2024-08-28
2024
-
[158]
Predicting the type and target of offensive posts in social media
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. Predicting the type and target of offensive posts in social media. arXiv preprint arXiv:1902.09666, 2019
1902 arXiv
-
[159]
Semeval-2020 task 12: Multilingual offensive language identification in social media (offenseval 2020).arXiv preprint arXiv:2006.07235, 2020
Marcos Zampieri, Preslav Nakov, Sara Rosenthal, Pepa Atanasova, Georgi Karadzhov, Hamdy Mubarak, Leon Derczynski, Zeses Pitenis, and C ¸ a˘ grı C ¸ ¨ oltekin. Semeval-2020 task 12: Multilingual offensive language identification in social media (offenseval 2020).arXiv preprint ...
2020 arXiv
-
[160]
Detecting hate speech on twitter using a convolution-gru based deep neural network
Ziqi Zhang, David Robinson, and Jonathan Tepper. Detecting hate speech on twitter using a convolution-gru based deep neural network. In The Semantic Web: 15th International Conference, ESWC 2018, Heraklion, Crete, Greece, June 3–7, 2018, Proceedings 15 , pages 745–760. Springe...
2018
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.