Pith. sign in

REVIEW 5 cited by

Hostility Detection Dataset in Hindi

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.03588 v1 pith:T2YGVOIB submitted 2020-11-06 cs.CL

classification cs.CL
keywords datasetdetectionhostilehostilitypostshindialongannotate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present a novel hostility detection dataset in Hindi language. We collect and manually annotate ~8200 online posts. The annotated dataset covers four hostility dimensions: fake news, hate speech, offensive, and defamation posts, along with a non-hostile label. The hostile posts are also considered for multi-label tags due to a significant overlap among the hostile classes. We release this dataset as part of the CONSTRAINT-2021 shared task on hostile post detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Web(er) of Hate: A Survey on How Hate Speech Is Typed

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Hate speech datasets vary because curators hold different ideal types of hate, so the field should document those assumptions instead of chasing a single definition.

  2. Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A contrastive audio-text framework detects hate speech in synthesized speech across six languages and outperforms baselines, with a new 127k-sample dataset.

  3. SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text Generation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    SMAB uses multi-armed bandit sampling and masked-language-model replacements to estimate word-level sensitivity of text classifiers, and applies it to accuracy prediction and adversarial text generation.

  4. HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A tri-partite survey finds that hate speech definitions and moderation practices in country laws, platform policies, and NLP datasets are largely misaligned, and calls for a unified proactive moderation framework.

  5. Explainable AI: XAI-Guided Context-Aware Data Augmentation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    XAI-guided augmentation that replaces the least important words, identified by Integrated Gradients, with back-translated synonyms or paraphrases improves hate speech and sentiment classification accuracy by up to 8 p...

Pith tools