Pith. sign in

REVIEW 1 cited by

Privacy Re-identification Attacks on Tabular GANs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00696 v1 pith:YEGLREU5 submitted 2024-03-31 cs.CR cs.LG

classification cs.CRcs.LG
keywords attackssamplessyntheticgenerativeprivacyre-identificationtrainingpotentially
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative models are subject to overfitting and thus may potentially leak sensitive information from the training data. In this work. we investigate the privacy risks that can potentially arise from the use of generative adversarial networks (GANs) for creating tabular synthetic datasets. For the purpose, we analyse the effects of re-identification attacks on synthetic data, i.e., attacks which aim at selecting samples that are predicted to correspond to memorised training samples based on their proximity to the nearest synthetic records. We thus consider multiple settings where different attackers might have different access levels or knowledge of the generative model and predictive, and assess which information is potentially most useful for launching more successful re-identification attacks. In doing so we also consider the situation for which re-identification attacks are formulated as reconstruction attacks, i.e., the situation where an attacker uses evolutionary multi-objective optimisation for perturbing synthetic samples closer to the training space. The results indicate that attackers can indeed pose major privacy risks by selecting synthetic samples that are likely representative of memorised training samples. In addition, we notice that privacy threats considerably increase when the attacker either has knowledge or has black-box access to the generative models. We also find that reconstruction attacks through multi-objective optimisation even increase the risk of identifying confidential samples.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Privacy-Utility Tradeoffs in Synthetic Smart Grid Data

    cs.LG 2025-05 conditional novelty 4.0 of 10

    On UK smart meter data, diffusion synthetic data yields the highest classifier accuracy (macro-F1 88.2%) while CTGAN full synthesis provides the strongest privacy protection against reconstruction attacks.

Pith tools