Pith. sign in

REVIEW 4 cited by

Applying sparse autoencoders to unlearn knowledge in language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19278 v2 pith:K72MCI25 submitted 2024-10-25 cs.LG cs.AI

classification cs.LGcs.AI
keywords featureslanguagemodelsunlearnautoencodersbiologyexistingknowledge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate whether sparse autoencoders (SAEs) can be used to remove knowledge from language models. We use the biology subset of the Weapons of Mass Destruction Proxy dataset and test on the gemma-2b-it and gemma-2-2b-it language models. We demonstrate that individual interpretable biology-related SAE features can be used to unlearn a subset of WMDP-Bio questions with minimal side-effects in domains other than biology. Our results suggest that negative scaling of feature activations is necessary and that zero ablating features is ineffective. We find that intervening using multiple SAE features simultaneously can unlearn multiple different topics, but with similar or larger unwanted side-effects than the existing Representation Misdirection for Unlearning technique. Current SAE quality or intervention techniques would need to improve to make SAE-based unlearning comparable to the existing fine-tuning based techniques.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Position: Use Sparse Autoencoders to Discover Unknowns

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Sparse autoencoders are best used to discover unknown concepts, not to act on known concepts.

  2. Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.

  3. Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

    cs.CL 2025-08 reject novelty 5.0 of 10

    Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.

  4. Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models

    cs.CL 2025-05 reject novelty 5.0 of 10

    Activation steering guided by verbal-versus-symbolic CoT decomposition improves math accuracy by a few points, but the SAE-free derivation conflates L1 and L2 objectives.

Pith tools