Pith. sign in

REVIEW

Developing a Multilingual Annotated Corpus of Misogyny and Aggression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.07428 v1 pith:WAO5QQDI submitted 2020-03-16 cs.CL

classification cs.CL
keywords misogynyaggressionannotatedcommentsaggressiveannotationcorpusdiscuss
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we discuss the development of a multilingual annotated corpus of misogyny and aggression in Indian English, Hindi, and Indian Bangla as part of a project on studying and automatically identifying misogyny and communalism on social media (the ComMA Project). The dataset is collected from comments on YouTube videos and currently contains a total of over 20,000 comments. The comments are annotated at two levels - aggression (overtly aggressive, covertly aggressive, and non-aggressive) and misogyny (gendered and non-gendered). We describe the process of data collection, the tagset used for annotation, and issues and challenges faced during the process of annotation. Finally, we discuss the results of the baseline experiments conducted to develop a classifier for misogyny in the three languages.

Discussion (0). Continue with ORCID to comment.

Pith tools