Pith. sign in

REVIEW 1 cited by

The Poison of Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.13449 v1 pith:3XGKVZ7N submitted 2023-08-25 cs.CL

classification cs.CL
keywords alignmentcontentdatasetinstructionlanguagemodelmodelsperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

From the perspective of content safety issues, alignment has shown to limit large language models' (LLMs) harmful content generation. This intentional method of reinforcing models to not respond to certain user inputs seem to be present in many modern open-source instruction tuning datasets such as OpenAssistant or Guanaco. We introduce a novel insight to an instruction-tuned model's performance affected by the presence of alignment in supervised fine-tuning dataset. To be specific, we noticed that alignment acts as if it is poisoning the instruction dataset. Experimentally, we demonstrate that aligned answers significantly worsen the performance of the resulting fine-tuned model's on various reasoning benchmarks such as Big Bench (BBH), Massive Multitask Language Understanding (MMLU), Human Eval, and Discrete Reasoning Over Paragraphs (DROP), performing worse than the counterpart tuned without alignment by 4-33%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Perspective Dial: Measuring Perspective of Text and Guiding LLM Outputs

    cs.CL 2025-06 reject novelty 5.0 of 10

    Perspective-Dial uses contrastive learning to build a perspective metric and greedy prompt optimization to steer LLM outputs toward a user-chosen viewpoint.

Pith tools