Pith. sign in

REVIEW 2 cited by

Feature Unlearning for Pre-trained GANs and VAEs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.05699 v4 pith:BKF2P3Q2 submitted 2023-03-10 cs.CV cs.LG

classification cs.CVcs.LG
keywords featurepre-trainedtargetunlearningimagemodelfeaturesexperiments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We tackle the problem of feature unlearning from a pre-trained image generative model: GANs and VAEs. Unlike a common unlearning task where an unlearning target is a subset of the training set, we aim to unlearn a specific feature, such as hairstyle from facial images, from the pre-trained generative models. As the target feature is only presented in a local region of an image, unlearning the entire image from the pre-trained model may result in losing other details in the remaining region of the image. To specify which features to unlearn, we collect randomly generated images that contain the target features. We then identify a latent representation corresponding to the target feature and then use the representation to fine-tune the pre-trained model. Through experiments on MNIST, CelebA, and FFHQ datasets, we show that target features are successfully removed while keeping the fidelity of the original models. Further experiments with an adversarial attack show that the unlearned model is more robust under the presence of malicious parties.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DECAF: De-Clustering for Adaptive Representational Unlearning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    DECAF is a forget-only unlearning method that adds input noise, suppresses the forget-class probability, and diversifies outputs, achieving 0.10% forget accuracy and 79.4% retain accuracy on CIFAR-10/ResNet-18 while d...

  2. Realistic Image-to-Image Machine Unlearning via Decoupling and Knowledge Retention

    cs.LG 2025-02 reject novelty 4.0 of 10

    The paper claims that gradient ascent on forget samples makes them out-of-distribution for an unlearned image-to-image model, with formal guarantees and a data-poisoning audit.

Pith tools