Pith. sign in

REVIEW 1 cited by

Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.08593 v1 pith:45CGL456 submitted 2025-06-10 cs.CL

classification cs.CL
keywords hatespeechannotationmodelstraitsdetectionhatefulhuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hate speech detection is a socially sensitive and inherently subjective task, with judgments often varying based on personal traits. While prior work has examined how socio-demographic factors influence annotation, the impact of personality traits on Large Language Models (LLMs) remains largely unexplored. In this paper, we present the first comprehensive study on the role of persona prompts in hate speech classification, focusing on MBTI-based traits. A human annotation survey confirms that MBTI dimensions significantly affect labeling behavior. Extending this to LLMs, we prompt four open-source models with MBTI personas and evaluate their outputs across three hate speech datasets. Our analysis uncovers substantial persona-driven variation, including inconsistencies with ground truth, inter-persona disagreement, and logit-level biases. These findings highlight the need to carefully define persona prompts in LLM-based annotation workflows, with implications for fairness and alignment with human values.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Monolingual to Bilingual: Investigating Language Conditioning in Large Language Models for Psycholinguistic Tasks

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Prompted language identity changes both the outputs and the internal layer representations of Llama-3.3-70B and Qwen2.5-72B on sound symbolism and word valence tasks.

Reference graph

Works this paper leans on

4 extracted references · 2 canonical work pages · cited by 1 Pith paper

  1. [4]

    All [religious group]s are terrorists. Deport them all!

    Large language models for data annotation and synthesis: A survey. In Proceedings of the 2024 Con- ference on Empirical Methods in Natural Language Processing, pages 930–957, Miami, Florida, USA. Association for Computational Linguistics. Alexandra N Uma, Tommaso Fornaciari, Dirk Hovy, Sil- viu Paun, Barbara Plank, and Massimo Poesio. 2021. Learning from ...

  2. [2021]

    HateCheck: Functional tests for hate speech detection models. In Proceedings of the 59th An- nual Meeting of the Association for Computational Linguistics and the 11th International Joint Confer- ence on Natural Language Processing (Volume 1: Long Papers), pages 41–58, Online. Association for Computational Linguistics. Maarten Sap, Swabha Swayamdipta, Lau...

  3. [2024]

    description of personality

    Can LLM be a personalized judge? In Find- ings of the Association for Computational Linguistics: EMNLP 2024, pages 10126–10141, Miami, Florida, USA. Association for Computational Linguistics. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, and et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Eve Fleisig, Rediet Abebe, and Dan Kle...

  4. [2025]

    In Proceedings of the 31st International Conference on Computational Linguis- tics, pages 5071–5081, Abu Dhabi, UAE

    Can large language models understand you bet- ter? an MBTI personality detection dataset aligned with population traits. In Proceedings of the 31st International Conference on Computational Linguis- tics, pages 5071–5081, Abu Dhabi, UAE. Associa- tion for Computational Linguistics. Andy Liu, Mona Diab, and Daniel Fried. 2024. Evalu- ating large language m...

Pith tools