Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth Interna- tional Conference on Learning Representations
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it