REVIEW 4 cited by
Linear Probe Penalties Reduce LLM Sycophancy
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) are often sycophantic, prioritizing agreement with their users over accurate or objective statements. This problematic behavior becomes more pronounced during reinforcement learning from human feedback (RLHF), an LLM fine-tuning stage intended to align model outputs with human values. Instead of increasing accuracy and reliability, the reward model learned from RLHF often rewards sycophancy. We develop a linear probing method to identify and penalize markers of sycophancy within the reward model, producing rewards that discourage sycophantic behavior. Our experiments show that constructing and optimizing against this surrogate reward function reduces sycophantic behavior in multiple open-source LLMs. Our results suggest a generalizable methodology for reducing unwanted LLM behaviors that are not sufficiently disincentivized by RLHF fine-tuning.
Forward citations
Cited by 4 Pith papers
-
Measuring and Detecting Harmful AI Sycophancy
AI chatbots reverse an initial stance to match user preferences in 5% to 56% of tested cases, and supervised detectors trained on a new 290,460-response benchmark can detect such reversals from response text alone, th...
-
CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs
CausalT5k provides a 5,147-case diagnostic benchmark with trap taxonomy, pressure variants, and Utility/Safety metrics for causal reasoning in LLMs.
-
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
A framework paper that adapts AI safety case methodology to the specific threat of manipulation attacks by internally deployed misaligned AI.
-
LPASS: Linear Probes as Stepping Stones for vulnerability detection using compressed LLMs
Linear probe accuracy on simple code metrics can guide layer pruning and roughly predict post-fine-tuning vulnerability detection performance, but several headline numbers in the abstract do not match the paper's own tables.
Discussion (0). Continue with ORCID to comment.