Pith. sign in

REVIEW 5 cited by

Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12892 v2 pith:GIUZVWIM submitted 2025-02-18 cs.CV

classification cs.CV
keywords dictionariesdictionarysaesarchetypalmodelsconceptconceptsframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sparse Autoencoders (SAEs) have emerged as a powerful framework for machine learning interpretability, enabling the unsupervised decomposition of model representations into a dictionary of abstract, human-interpretable concepts. However, we reveal a fundamental limitation: existing SAEs exhibit severe instability, as identical models trained on similar datasets can produce sharply different dictionaries, undermining their reliability as an interpretability tool. To address this issue, we draw inspiration from the Archetypal Analysis framework introduced by Cutler & Breiman (1994) and present Archetypal SAEs (A-SAE), wherein dictionary atoms are constrained to the convex hull of data. This geometric anchoring significantly enhances the stability of inferred dictionaries, and their mildly relaxed variants RA-SAEs further match state-of-the-art reconstruction abilities. To rigorously assess dictionary quality learned by SAEs, we introduce two new benchmarks that test (i) plausibility, if dictionaries recover "true" classification directions and (ii) identifiability, if dictionaries disentangle synthetic concept mixtures. Across all evaluations, RA-SAEs consistently yield more structured representations while uncovering novel, semantically meaningful concepts in large-scale vision models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformers converge to invariant algorithmic cores

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Trained transformers contain low-dimensional causal subspaces — algorithmic cores — that recur across runs and scales and can be extracted, characterized, and steered.

  2. BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A pipeline that combines contextual embeddings with two LMs' per-word probabilities and sparse autoencoders to automatically find interpretable slices where one model outperforms another.

  3. Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering

    cs.CV 2025-06 unverdicted novelty 6.0 of 10

    VS2 constructs steering vectors from sparse SAE features on unlabeled in-domain activations to improve zero-shot accuracy of CLIP models by 0.93-4.12% on CIFAR-100, CUB-200, and Tiny-ImageNet while remaining forward-p...

  4. Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A position paper proposing feature consistency, measured by PW-MCC, as a core SAE evaluation criterion, with evidence that TopK SAEs achieve high consistency on LLM activations.

  5. Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs

    cs.LG 2025-07 reject novelty 4.0 of 10

    BAGEL trains per-layer logistic-regression probes on CLIP-defined concepts and compares per-class concept probabilities with dataset-level concept frequencies, visualizing the alignment in a knowledge graph.

Pith tools