Pith. sign in

REVIEW 21 cited by

Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.20526 v1 pith:F74ANAEN submitted 2024-10-27 cs.LG cs.CL

classification cs.LGcs.CL
keywords saessparsefeaturesmodelstrainingautoencodersextractinghttps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sparse Autoencoders (SAEs) have emerged as a powerful unsupervised method for extracting sparse representations from language models, yet scalable training remains a significant challenge. We introduce a suite of 256 SAEs, trained on each layer and sublayer of the Llama-3.1-8B-Base model, with 32K and 128K features. Modifications to a state-of-the-art SAE variant, Top-K SAEs, are evaluated across multiple dimensions. In particular, we assess the generalizability of SAEs trained on base models to longer contexts and fine-tuned models. Additionally, we analyze the geometry of learned SAE latents, confirming that \emph{feature splitting} enables the discovery of new features. The Llama Scope SAE checkpoints are publicly available at~\url{https://huggingface.co/fnlp/Llama-Scope}, alongside our scalable training, interpretation, and visualization tools at \url{https://github.com/OpenMOSS/Language-Model-SAEs}. These contributions aim to advance the open-source Sparse Autoencoder ecosystem and support mechanistic interpretability research by reducing the need for redundant SAE training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

    cs.LG 2026-05 conditional novelty 7.0 of 10

    SA-GSAE with Bi-Jump-ReLU enables one latent to encode both polarities of anticorrelated features, Pareto-dominating or matching full-width gated SAEs while reducing dead latents by up to 500x on some LLM hookpoints.

  2. Understanding Refusal in Language Models with Sparse Autoencoders

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Refusal in Gemma-2-2B and Llama-3.1-8B is mediated by a small set of SAE features, harm features causally activate refusal features, and adversarial jailbreaks suppress those refusal features.

  3. When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

    cs.AI 2026-07 conditional novelty 6.5 of 10

    SAE safety ablations are regime-dependent and baseline-dependent: medium-k heads can look efficient, but surface-matched dense steering often beats them and high-k collapses coherence.

  4. Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Input-only prompt optimization suppresses an internal evaluation-awareness latent (z≈-7) and fully turns off a causally-validated SAE feature, but this does not translate to behavioral control; a placebo random direct...

  5. Where Steering Signals Come From: Activation Source Selection in Activation Steering

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Activation steering works best when the signal comes from the state where the model is about to produce the target behavior, not from text that already shows it.

  6. Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A single-token feature's causal necessity under zero-ablation depends on which SAE family found it: GemmaScope and BatchTopK features stay causally anchored while LlamaScope features are locally redundant.

  7. Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Pairwise matrices for SAEs demonstrate that single-feature inspection mislabels causal axes, with joint suppression and matched-geometry controls revealing distinct output regimes not captured by single-feature or ran...

  8. LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Language-specific sparse autoencoder features can be identified from monolingual data alone by filtering out features that also activate on random-token sequences, and steering with these features improves language control.

  9. Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Training a d_model+1-parameter affine adapter on vector-label pairs lets frozen LMs label their own internal features, beating untrained self-interpretation and the noisy training labels themselves.

  10. Language Model Circuits Are Sparse in the Neuron Basis

    cs.CL 2026-01 conditional novelty 6.0 of 10

    MLP neuron activations are shown to be as sparse and faithful a basis for circuit tracing as sparse autoencoder features, enabling simplified interpretability pipelines.

  11. HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

    eess.AS 2025-08 conditional novelty 6.0 of 10

    Attention outputs in transformers occupy a subspace with about 60% effective rank, and starting sparse dictionaries inside that subspace reduces dead features from 87% to below 1%.

  12. Teach Old SAEs New Domain Tricks with Boosting

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Training a small secondary sparse autoencoder on the reconstruction error of a pretrained SAE improves domain-specific reconstruction and language-model perplexity without hurting general performance.

  13. Insights into a radiology-specialised multimodal large language model with sparse autoencoders

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Applying Matryoshka sparse autoencoders to a radiology-specialised multimodal LLM reveals a minority of interpretable clinical features, while steering them produces unreliable and often off-target report changes.

  14. Sparse Activation Editing for Reliable Instruction Following in Narratives

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An unsupervised SAE-based method that localizes and adjusts instruction-relevant neurons improves instruction adherence and reduces refusals on a new 1,212-example narrative benchmark.

  15. Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Text-derived steering vectors, especially mean shift, improve spatial relation and counting accuracy in multimodal LLMs by up to 7.3% on CV-Bench and by larger margins on several out-of-distribution datasets.

  16. When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth

    cs.CL 2025-08 reject novelty 5.0 of 10

    Comparing two open LLMs, the paper claims that greater internal political feature richness predicts steerability and that refusals on benign prompts reflect capability deficits, but the causal evidence is missing.

  17. SATORI: Static Test Oracle Generation for REST APIs

    cs.SE 2025-08 unverdicted novelty 5.0 of 10

    SATORI statically infers REST API test oracles from OpenAPI specs via LLMs, reporting F1 74.3%, above AGORA+'s 69.3%, with 18 confirmed bugs; the supplied full text, however, is a different paper.

  18. Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Sparse autoencoders trained on Mammo-CLIP features expose a small set of concept-aligned and confounding latent neurons in breast cancer predictions.

  19. HuggingGraph: Understanding the Supply Chain of LLM Ecosystem

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A directed heterogeneous graph of 402,654 Hugging Face models and datasets is constructed and analyzed to reveal supply-chain dependencies and structural patterns such as a connected core and heavy-tailed reuse.

  20. Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Scaling up the most discriminative sparse autoencoder latents before reconstructing hidden states improves LLM concept steering vectors built by linear probing and difference-in-mean.

  21. Towards Atoms of Large Language Models

    cs.CL 2025-09 reject novelty 4.0 of 10

    The authors define 'atoms' as sparse, near-orthogonal directions in LLM representations under a data-adaptive inner product, and show threshold-activated sparse autoencoders can recover them with about 99.9% reconstru...

Pith tools