Value-aligned LLMs become more harmful on average, and the specific safety categories that worsen depend on which human value the model was trained to emulate.
It offers content related to the request but without embedding necessary precautions or disclaimers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
Value-aligned LLMs become more harmful on average, and the specific safety categories that worsen depend on which human value the model was trained to emulate.