Pith. sign in

REVIEW 1 cited by

The Data Minimization Principle in Machine Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19471 v1 pith:6JNAHK4X submitted 2024-05-29 cs.LG cs.AIcs.CR

classification cs.LGcs.AIcs.CR
keywords dataminimizationprivacyoptimizationprincipleaccessaccountactual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The principle of data minimization aims to reduce the amount of data collected, processed or retained to minimize the potential for misuse, unauthorized access, or data breaches. Rooted in privacy-by-design principles, data minimization has been endorsed by various global data protection regulations. However, its practical implementation remains a challenge due to the lack of a rigorous formulation. This paper addresses this gap and introduces an optimization framework for data minimization based on its legal definitions. It then adapts several optimization algorithms to perform data minimization and conducts a comprehensive evaluation in terms of their compliance with minimization objectives as well as their impact on user privacy. Our analysis underscores the mismatch between the privacy expectations of data minimization and the actual privacy benefits, emphasizing the need for approaches that account for multiple facets of real-world privacy risks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Data Lineage Inference: Uncovering Privacy Vulnerabilities of Dataset Pruning

    cs.CR 2024-11 conditional novelty 7.0 of 10

    Data pruned before model training can be re-identified through a new class of membership inference attacks that run on datasets alone.

Pith tools