Pith. sign in

REVIEW 2 cited by

Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.08035 v1 pith:3THV5HKG submitted 2024-03-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmshatespeechhatefullanguageanalysischallengesclassifying
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) excel in many diverse applications beyond language generation, e.g., translation, summarization, and sentiment analysis. One intriguing application is in text classification. This becomes pertinent in the realm of identifying hateful or toxic speech -- a domain fraught with challenges and ethical dilemmas. In our study, we have two objectives: firstly, to offer a literature review revolving around LLMs as classifiers, emphasizing their role in detecting and classifying hateful or toxic content. Subsequently, we explore the efficacy of several LLMs in classifying hate speech: identifying which LLMs excel in this task as well as their underlying attributes and training. Providing insight into the factors that contribute to an LLM proficiency (or lack thereof) in discerning hateful content. By combining a comprehensive literature review with an empirical analysis, our paper strives to shed light on the capabilities and constraints of LLMs in the crucial domain of hate speech detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extracting Participation in Collective Action from Social Media

    cs.SI 2025-01 conditional novelty 6.0 of 10

    A new classifier suite detects and levels expressions of collective action participation in Reddit comments, reaching weighted F1=0.71 for binary detection.

  2. Towards Efficient and Explainable Hate Speech Detection via Model Distillation

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A distilled Llama-3-8B model, trained with chain-of-thought rationales from Llama-3-70B, matches the large model's explanation quality and shows higher F1 on a hate speech test sample.

Pith tools