REVIEW 4 cited by
Evaluations of Machine Learning Privacy Defenses are Misleading
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Empirical defenses for machine learning privacy forgo the provable guarantees of differential privacy in the hope of achieving higher utility while resisting realistic adversaries. We identify severe pitfalls in existing empirical privacy evaluations (based on membership inference attacks) that result in misleading conclusions. In particular, we show that prior evaluations fail to characterize the privacy leakage of the most vulnerable samples, use weak attacks, and avoid comparisons with practical differential privacy baselines. In 5 case studies of empirical privacy defenses, we find that prior evaluations underestimate privacy leakage by an order of magnitude. Under our stronger evaluation, none of the empirical defenses we study are competitive with a properly tuned, high-utility DP-SGD baseline (with vacuous provable guarantees).
Forward citations
Cited by 4 Pith papers
-
SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation
A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending ...
-
Granite Guardian
Granite Guardian 2B and 8B are open-source LLM guardrails that detect harmful content, jailbreaks, and RAG hallucination risks, reporting AUC 0.871 on harm benchmarks and 0.854 on groundedness benchmarks.
-
One-shot Federated Learning via Synthetic Distiller-Distillate Communication
FedSD2C beats prior one-shot federated learning baselines on ImageNette, Tiny-ImageNet, and OpenImage by sending compact latent codes of selected, Fourier-perturbed images instead of local models.
-
Underestimated Privacy Risks for Minority Populations in Large Language Model Unlearning
Standard random-data LLM unlearning evaluations understate privacy leakage for minority data, as shown by canary and real rare-PII experiments across three datasets and two models.
Discussion (0). Continue with ORCID to comment.