REVIEW 1 cited by
K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment
read the original abstract
Online hate speech detection has become an important issue due to the growth of online content, but resources in languages other than English are extremely limited. We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns. The dataset consists of 109k utterances from news comments and provides a multi-label classification using 1 to 4 labels, and handles subjectivity and intersectionality. We evaluate strong baseline experiments on K-MHaS using Korean-BERT-based language models with six different metrics. KR-BERT with a sub-character tokenizer outperforms others, recognizing decomposed characters in each hate speech class.
Forward citations
Cited by 1 Pith paper
-
AEGIS: Awareness-Enhanced Guidance for Iterative Safeguard
Marking offensive spans changes — but does not consistently improve — the toxicity–meaning trade-off in multilingual detoxification; the effect depends on the generator backbone and the language.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.