Pith. sign in

REVIEW 2 cited by

An Overview of Privacy in Machine Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.08679 v1 pith:GHQG5D75 submitted 2020-05-18 cs.LG cs.AIcs.CRcs.CYstat.ML

classification cs.LGcs.AIcs.CRcs.CYstat.ML
keywords informationlearningmachinedatamodelsprivacyreviewaccess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Over the past few years, providers such as Google, Microsoft, and Amazon have started to provide customers with access to software interfaces allowing them to easily embed machine learning tasks into their applications. Overall, organizations can now use Machine Learning as a Service (MLaaS) engines to outsource complex tasks, e.g., training classifiers, performing predictions, clustering, etc. They can also let others query models trained on their data. Naturally, this approach can also be used (and is often advocated) in other contexts, including government collaborations, citizen science projects, and business-to-business partnerships. However, if malicious users were able to recover data used to train these models, the resulting information leakage would create serious issues. Likewise, if the inner parameters of the model are considered proprietary information, then access to the model should not allow an adversary to learn such parameters. In this document, we set to review privacy challenges in this space, providing a systematic review of the relevant research literature, also exploring possible countermeasures. More specifically, we provide ample background information on relevant concepts around machine learning and privacy. Then, we discuss possible adversarial models and settings, cover a wide range of attacks that relate to private and/or sensitive information leakage, and review recent results attempting to defend against such attacks. Finally, we conclude with a list of open problems that require more work, including the need for better evaluations, more targeted defenses, and the study of the relation to policy and data protection efforts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Federated Learning for Commercial Image Sources

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The authors present a new 31-class, 8-source image classification dataset for federated learning and show that Fed-Cyclic and Fed-Star beat FedAvg and RingFed on it.

  2. When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning

    cs.CR 2025-06 conditional novelty 5.0 of 10

    Membership inference against contrastive encoders is more accurate for larger frameworks and backbones, and a lightweight likelihood attack based on feature-vector p-norms matches or beats prior attacks with fewer queries.

Pith tools