Pith. sign in

REVIEW 4 cited by

Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01218 v3 pith:TIH6RDWW submitted 2024-03-02 cs.LG cs.CR

classification cs.LGcs.CR
keywords unlearningexamplesmodelu-miasexampleprivacytrainingattacker
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The high cost of model training makes it increasingly desirable to develop techniques for unlearning. These techniques seek to remove the influence of a training example without having to retrain the model from scratch. Intuitively, once a model has unlearned, an adversary that interacts with the model should no longer be able to tell whether the unlearned example was included in the model's training set or not. In the privacy literature, this is known as membership inference. In this work, we discuss adaptations of Membership Inference Attacks (MIAs) to the setting of unlearning (leading to their "U-MIA" counterparts). We propose a categorization of existing U-MIAs into "population U-MIAs", where the same attacker is instantiated for all examples, and "per-example U-MIAs", where a dedicated attacker is instantiated for each example. We show that the latter category, wherein the attacker tailors its membership prediction to each example under attack, is significantly stronger. Indeed, our results show that the commonly used U-MIAs in the unlearning literature overestimate the privacy protection afforded by existing unlearning techniques on both vision and language models. Our investigation reveals a large variance in the vulnerability of different examples to per-example U-MIAs. In fact, several unlearning algorithms lead to a reduced vulnerability for some, but not all, examples that we wish to unlearn, at the expense of increasing it for other examples. Notably, we find that the privacy protection for the remaining training examples may worsen as a consequence of unlearning. We also discuss the fundamental difficulty of equally protecting all examples using existing unlearning schemes, due to the different rates at which examples are unlearned. We demonstrate that naive attempts at tailoring unlearning stopping criteria to different examples fail to alleviate these issues.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Gradient Concentration, Not Weight Saliency, Explains Representation-Level Class Unlearning

    cs.LG 2026-07 accept novelty 7.0 of 10

    In class unlearning on CIFAR-10/100 with ResNet-18, the identity of saliency-selected weights does not affect representation-level recovery; late-layer gradient concentration and representation geometry drive the outcome.

  2. System-Aware Unlearning Algorithms: Use Lesser, Forget Faster

    cs.LG 2025-06 conditional novelty 7.0 of 10

    The paper introduces system-aware unlearning and gives the first exact unlearning algorithm for linear classification that stores a sublinear-size core set instead of the entire dataset.

  3. On the Necessity of Output Distribution Reweighting for Effective Class Unlearning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Class unlearning methods leak membership through neighbor-class output probabilities, and a tilted reweighting objective that mimics retrained models reduces this leakage.

  4. Leveraging Per-Instance Privacy for Machine Unlearning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Per-instance privacy losses, estimated from gradient norms during training, predict the number of fine-tuning steps needed for machine unlearning and rank data points by unlearning difficulty.

Pith tools