REVIEW 1 cited by
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Hate speech detection is a socially sensitive and inherently subjective task, with judgments often varying based on personal traits. While prior work has examined how socio-demographic factors influence annotation, the impact of personality traits on Large Language Models (LLMs) remains largely unexplored. In this paper, we present the first comprehensive study on the role of persona prompts in hate speech classification, focusing on MBTI-based traits. A human annotation survey confirms that MBTI dimensions significantly affect labeling behavior. Extending this to LLMs, we prompt four open-source models with MBTI personas and evaluate their outputs across three hate speech datasets. Our analysis uncovers substantial persona-driven variation, including inconsistencies with ground truth, inter-persona disagreement, and logit-level biases. These findings highlight the need to carefully define persona prompts in LLM-based annotation workflows, with implications for fairness and alignment with human values.
Forward citations
Cited by 1 Pith paper
-
From Monolingual to Bilingual: Investigating Language Conditioning in Large Language Models for Psycholinguistic Tasks
Prompted language identity changes both the outputs and the internal layer representations of Llama-3.3-70B and Qwen2.5-72B on sound symbolism and word valence tasks.
Reference graph
Works this paper leans on
-
[4]
All [religious group]s are terrorists. Deport them all!
Large language models for data annotation and synthesis: A survey. In Proceedings of the 2024 Con- ference on Empirical Methods in Natural Language Processing, pages 930–957, Miami, Florida, USA. Association for Computational Linguistics. Alexandra N Uma, Tommaso Fornaciari, Dirk Hovy, Sil- viu Paun, Barbara Plank, and Massimo Poesio. 2021. Learning from ...
arXiv 2024
-
[2021]
HateCheck: Functional tests for hate speech detection models. In Proceedings of the 59th An- nual Meeting of the Association for Computational Linguistics and the 11th International Joint Confer- ence on Natural Language Processing (Volume 1: Long Papers), pages 41–58, Online. Association for Computational Linguistics. Maarten Sap, Swabha Swayamdipta, Lau...
work page 2022
-
[2024]
Can LLM be a personalized judge? In Find- ings of the Association for Computational Linguistics: EMNLP 2024, pages 10126–10141, Miami, Florida, USA. Association for Computational Linguistics. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, and et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Eve Fleisig, Rediet Abebe, and Dan Kle...
arXiv 2024
-
[2025]
Can large language models understand you bet- ter? an MBTI personality detection dataset aligned with population traits. In Proceedings of the 31st International Conference on Computational Linguis- tics, pages 5071–5081, Abu Dhabi, UAE. Associa- tion for Computational Linguistics. Andy Liu, Mona Diab, and Daniel Fried. 2024. Evalu- ating large language m...
work page 2024
Discussion (0). Sign in to comment.