Pith. sign in

REVIEW 1 cited by

HateModerate: Testing Hate Speech Detectors against Content Moderation Policies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.12418 v2 pith:YCS3XQJC submitted 2023-07-23 cs.SE cs.CL

classification cs.SEcs.CL
keywords contenthatepoliciesspeechhatemoderateautomateddetectorsexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To protect users from massive hateful content, existing works studied automated hate speech detection. Despite the existing efforts, one question remains: do automated hate speech detectors conform to social media content policies? A platform's content policies are a checklist of content moderated by the social media platform. Because content moderation rules are often uniquely defined, existing hate speech datasets cannot directly answer this question. This work seeks to answer this question by creating HateModerate, a dataset for testing the behaviors of automated content moderators against content policies. First, we engage 28 annotators and GPT in a six-step annotation process, resulting in a list of hateful and non-hateful test suites matching each of Facebook's 41 hate speech policies. Second, we test the performance of state-of-the-art hate speech detectors against HateModerate, revealing substantial failures these models have in their conformity to the policies. Third, using HateModerate, we augment the training data of a top-downloaded hate detector on HuggingFace. We observe significant improvement in the models' conformity to content policies while having comparable scores on the original test data. Our dataset and code can be found in the attachment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ModelCitizens: Representing Community Voices in Online Safety

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A community-annotated toxicity dataset with conversational context shows that models trained on ingroup labels outperform state-of-the-art moderation APIs.

Pith tools