REVIEW 4 cited by
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-truth collection of interpretable features that we expect good SAEs to recover. We thus propose to measure progress in interpretable dictionary learning by working in the setting of LMs trained on chess and Othello transcripts. These settings carry natural collections of interpretable features -- for example, "there is a knight on F3" -- which we leverage into $\textit{supervised}$ metrics for SAE quality. To guide progress in interpretable dictionary learning, we introduce a new SAE training technique, $\textit{p-annealing}$, which improves performance on prior unsupervised metrics as well as our new metrics.
Forward citations
Cited by 4 Pith papers
-
Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations
SA-GSAE with Bi-Jump-ReLU enables one latent to encode both polarities of anticorrelated features, Pareto-dominating or matching full-width gated SAEs while reducing dead latents by up to 500x on some LLM hookpoints.
-
Transformers Use Causal World Models in Maze-Solving Tasks
Maze-solving transformers store a causal, steerable map of maze connections in sparse features, and activating these features is more effective than removing them.
-
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
SHIFT-based and TPP metrics grade sparse autoencoders by whether their features can isolate and erase one concept without damaging others, and the paper shows they separate SAE architectures and improve during training.
-
InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders
Sparse autoencoders trained on ESM-2 recover thousands of interpretable features that align with Swiss-Prot concepts and can influence sequence generation.
Discussion (0). Continue with ORCID to comment.