Pith. sign in

REVIEW 3 cited by

Disentangling Dense Embeddings with Sparse Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.00657 v2 pith:VFH6I4MB submitted 2024-08-01 cs.LG

classification cs.LG
keywords embeddingssparsefeaturessemanticautoencodersdensesaesconcepts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sparse autoencoders (SAEs) have shown promise in extracting interpretable features from complex neural networks. We present one of the first applications of SAEs to dense text embeddings from large language models, demonstrating their effectiveness in disentangling semantic concepts. By training SAEs on embeddings of over 420,000 scientific paper abstracts from computer science and astronomy, we show that the resulting sparse representations maintain semantic fidelity while offering interpretability. We analyse these learned features, exploring their behaviour across different model capacities and introducing a novel method for identifying ``feature families'' that represent related concepts at varying levels of abstraction. To demonstrate the practical utility of our approach, we show how these interpretable features can be used to precisely steer semantic search, allowing for fine-grained control over query semantics. This work bridges the gap between the semantic richness of dense embeddings and the interpretability of sparse representations. We open source our embeddings, trained sparse autoencoders, and interpreted features, as well as a web app for exploring them.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sparse Autoencoders, Again?

    cs.LG 2025-06 conditional novelty 6.0 of 10

    VAEase gates the VAE decoder input by the encoder's variance, combining sparse-autoencoder adaptive sparsity with a hyperparameter-free loss; a global-minimizer theorem says active latent dimensions recover per-manifo...

  2. Isotropy Cliffs: The Geometric Signature of Decision-Making in Large Language Models

    cs.AI 2026-08 reject novelty 5.0 of 10

    Across five LLMs, a sharp decrease in embedding isotropy at a critical layer predicts multiple-choice accuracy, with Spearman correlations up to -0.92.

  3. Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning

    cs.CL 2025-06 conditional novelty 5.0 of 10

    State-of-the-art text embeddings lag far behind on tasks requiring pragmatic inference, stance detection, and social meaning, relative to their strong performance on surface semantic benchmarks.

Pith tools