In p-IgGen, TopK SAE features correlate with IGHJ4 but do not steer it, while two Ordered SAE latents predictably raise and lower IGHJ4 output.
(19) Bechtel, W.; Richardson, R
2 Pith papers cite this work, alongside 31 external citations. Polarity classification is still indexing.
2
Pith papers citing it
31
external citations · OpenAlex
fields
cs.LG 2representative citing papers
citing papers explorer
-
Mechanistic Interpretability of Antibody Language Models Using SAEs
In p-IgGen, TopK SAE features correlate with IGHJ4 but do not steer it, while two Ordered SAE latents predictably raise and lower IGHJ4 output.
- Enhancing AI Interpretability with Localised Architectures