Pith. sign in

REVIEW 4 cited by

Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.00113 v2 pith:AQW4SOVV submitted 2024-07-31 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords interpretablefeaturesdictionarylearningmetricsprogresssaeslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-truth collection of interpretable features that we expect good SAEs to recover. We thus propose to measure progress in interpretable dictionary learning by working in the setting of LMs trained on chess and Othello transcripts. These settings carry natural collections of interpretable features -- for example, "there is a knight on F3" -- which we leverage into $\textit{supervised}$ metrics for SAE quality. To guide progress in interpretable dictionary learning, we introduce a new SAE training technique, $\textit{p-annealing}$, which improves performance on prior unsupervised metrics as well as our new metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

    cs.LG 2026-05 conditional novelty 7.0 of 10

    SA-GSAE with Bi-Jump-ReLU enables one latent to encode both polarities of anticorrelated features, Pareto-dominating or matching full-width gated SAEs while reducing dead latents by up to 500x on some LLM hookpoints.

  2. Transformers Use Causal World Models in Maze-Solving Tasks

    cs.LG 2024-12 conditional novelty 7.0 of 10

    Maze-solving transformers store a causal, steerable map of maze connections in sparse features, and activating these features is more effective than removing them.

  3. Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks

    cs.LG 2024-11 conditional novelty 6.0 of 10

    SHIFT-based and TPP metrics grade sparse autoencoders by whether their features can isolate and erase one concept without damaging others, and the paper shows they separate SAE architectures and improve during training.

  4. InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders

    q-bio.BM 2024-11 conditional novelty 6.0 of 10

    Sparse autoencoders trained on ESM-2 recover thousands of interpretable features that align with Swiss-Prot concepts and can influence sequence generation.

Pith tools