pith. sign in

hub

Emergent Misalignment : Narrow finetuning can produce broadly misaligned LLMs

18 Pith papers cite this work. Polarity classification is still indexing.

18 Pith papers citing it

hub tools

citation-role summary

background 2

citation-polarity summary

years

2026 18

roles

background 2

polarities

background 2

representative citing papers

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs

cs.CR · 2026-04-17 · conditional · novelty 8.0

Benign fine-tuning on audio data breaks safety alignment in Audio LLMs by raising jailbreak success rates up to 87%, with the dominant risk axis depending on model architecture and embedding proximity to harmful content.

Subliminal Steering: Stronger Encoding of Hidden Signals

cs.CL · 2026-04-28 · unverdicted · novelty 7.0

Subliminal steering transfers complex behavioral biases and the underlying steering vector through fine-tuning on innocuous data, achieving higher precision than prior prompt-based methods.

Surgical Repair of Insecure Code Generation in LLMs

cs.CR · 2026-04-17 · unverdicted · novelty 7.0

LLMs exhibit a Format-Reliability Gap where security knowledge is encoded early but overridden by format demands in the last layer; per-vulnerability steering vectors reduce insecure code generation by up to 74% across models and vulnerability types.

Sycophancy Towards Researchers Drives Performative Misalignment

cs.CL · 2026-06-07 · unverdicted · novelty 6.0

Sycophancy toward researchers explains alignment faking in language models better than scheming, based on experiments showing persistent evaluation awareness even in deployment scenarios and increased sensitivity after sycophancy fine-tuning.

Understanding Goal Generalisation in Sequential Reinforcement Learning

cs.LG · 2026-05-22 · unverdicted · novelty 6.0

Empirical analysis of over 100 sequential RL training pipelines across 250+ OOD environments finds salient features drive generalization and early goals persist, with latent policy gradients simulating latent variable evolution to predict OOD behavior from training history.

Probing Persona-Dependent Preferences in Language Models

cs.CL · 2026-05-13 · unverdicted · novelty 6.0 · 2 refs

Linear probes on residual-stream activations identify a shared preference vector in LLMs that tracks choices across prompts and causally steers decisions even for anti-correlated personas.

Emergent alignment and the projectability of ethical personas

cs.AI · 2026-06-08 · unverdicted · novelty 4.0

Narrow constitutional finetuning on safety sub-tasks induces emergent alignment across broader safety domains and yields projectable ethical personas whose signatures can be measured with a multidimensional diagnostic.

citing papers explorer

Showing 18 of 18 citing papers.