Pith. sign in

REVIEW 3 cited by

Provably Unlearnable Data Examples

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.03316 v2 pith:NGEAZF5L submitted 2024-05-06 cs.LG cs.CR

classification cs.LGcs.CR
keywords dataunlearnablelearnabilityexamplesunauthorizedaccuracyattackerscertified
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The exploitation of publicly accessible data has led to escalating concerns regarding data privacy and intellectual property (IP) breaches in the age of artificial intelligence. To safeguard both data privacy and IP-related domain knowledge, efforts have been undertaken to render shared data unlearnable for unauthorized models in the wild. Existing methods apply empirically optimized perturbations to the data in the hope of disrupting the correlation between the inputs and the corresponding labels such that the data samples are converted into Unlearnable Examples (UEs). Nevertheless, the absence of mechanisms to verify the robustness of UEs against uncertainty in unauthorized models and their training procedures engenders several under-explored challenges. First, it is hard to quantify the unlearnability of UEs against unauthorized adversaries from different runs of training, leaving the soundness of the defense in obscurity. Particularly, as a prevailing evaluation metric, empirical test accuracy faces generalization errors and may not plausibly represent the quality of UEs. This also leaves room for attackers, as there is no rigid guarantee of the maximal test accuracy achievable by attackers. Furthermore, we find that a simple recovery attack can restore the clean-task performance of the classifiers trained on UEs by slightly perturbing the learned weights. To mitigate the aforementioned problems, in this paper, we propose a mechanism for certifying the so-called $(q, \eta)$-Learnability of an unlearnable dataset via parametric smoothing. A lower certified $(q, \eta)$-Learnability indicates a more robust and effective protection over the dataset. Concretely, we 1) improve the tightness of certified $(q, \eta)$-Learnability and 2) design Provably Unlearnable Examples (PUEs) which have reduced $(q, \eta)$-Learnability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. T2UE: Generating Unlearnable Examples from Text Descriptions

    cs.AI 2025-08 conditional novelty 7.0 of 10

    T2UE trains a text-to-noise generator with a frozen CLIP model so that noise derived from a caption can make any image unlearnable for later CLIP or supervised training.

  2. New presentation of the twisted Yangian of type $D$

    math.QA 2025-07 conditional novelty 7.0 of 10

    A new finite presentation of the twisted Yangian of type D is proven, and a conjectural twisted affine Yangian of type D is defined.

  3. DiffUE: Enhancing Utility-Unlearnability Trade-off of Unlearnable Examples via Diffusion Autoencoders

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Injecting defensive noise into the semantic latent of a diffusion autoencoder produces unlearnable images with superior quality-unlearnability trade-off and robustness to relearning attacks versus pixel-space baselines.

Pith tools