Pith. sign in

REVIEW 4 cited by

Data Bootstrapping Approaches to Improve Low Resource Abusive Language Detection for Indic Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.12543 v1 pith:PJFQ4LTF submitted 2022-04-26 cs.CL cs.LG

classification cs.CLcs.LG
keywords abusivespeechlanguagesmodelsdetectionindiclanguagevarious
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Abusive language is a growing concern in many social media platforms. Repeated exposure to abusive speech has created physiological effects on the target users. Thus, the problem of abusive language should be addressed in all forms for online peace and safety. While extensive research exists in abusive speech detection, most studies focus on English. Recently, many smearing incidents have occurred in India, which provoked diverse forms of abusive speech in online space in various languages based on the geographic location. Therefore it is essential to deal with such malicious content. In this paper, to bridge the gap, we demonstrate a large-scale analysis of multilingual abusive speech in Indic languages. We examine different interlingual transfer mechanisms and observe the performance of various multilingual models for abusive speech detection for eight different Indic languages. We also experiment to show how robust these models are on adversarial attacks. Finally, we conduct an in-depth error analysis by looking into the models' misclassified posts across various settings. We have made our code and models public for other researchers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter

    cs.CL 2024-11 conditional novelty 7.0 of 10

    HateDay provides a representative global sample of one day of Twitter and shows that hate speech detection models achieve much lower average precision on real-world data than on academic datasets.

  2. Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

    cs.CL 2026-07 conditional novelty 5.0 of 10

    A gated fusion head that conditions English toxicity, Indic abuse, and rule-based severity scores on the text context improves code-mixed abuse detection in 10/12 in-domain and 7/8 transfer comparisons.

  3. Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    The best overall F1 scores come from larger models such as Qwen2.5-VL and DeBERTa-v3, but performance is dataset-dependent and victim detection remains poor across all model families.

  4. NLPineers@ NLU of Devanagari Script Languages 2025: Hate Speech Detection using Ensembling of BERT-based models

    cs.CL 2024-12 conditional novelty 4.0 of 10

    An ensemble of XLM-RoBERTa, MuRIL, and an abusive-tuned MuRIL achieved 0.7762 recall and 0.6914 F1 on Devanagari hate speech detection.

Pith tools