Pith. sign in

REVIEW 1 cited by

TabularMark: Watermarking Tabular Datasets for Machine Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14841 v1 pith:HNQ7SXU4 submitted 2024-06-21 cs.CR cs.DBcs.LG

classification cs.CRcs.DBcs.LG
keywords datadatasetsutilitywatermarkingmodelstabulartabularmarkwhile
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Watermarking is broadly utilized to protect ownership of shared data while preserving data utility. However, existing watermarking methods for tabular datasets fall short on the desired properties (detectability, non-intrusiveness, and robustness) and only preserve data utility from the perspective of data statistics, ignoring the performance of downstream ML models trained on the datasets. Can we watermark tabular datasets without significantly compromising their utility for training ML models while preventing attackers from training usable ML models on attacked datasets? In this paper, we propose a hypothesis testing-based watermarking scheme, TabularMark. Data noise partitioning is utilized for data perturbation during embedding, which is adaptable for numerical and categorical attributes while preserving the data utility. For detection, a custom-threshold one proportion z-test is employed, which can reliably determine the presence of the watermark. Experiments on real-world and synthetic datasets demonstrate the superiority of TabularMark in detectability, non-intrusiveness, and robustness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Watermarking Generative Categorical Data

    cs.CR 2024-11 conditional novelty 5.0 of 10

    A distribution-level watermark for categorical data, embedded by secret hashing and detected by inverting the hash mixture and comparing total variation distance to the original distribution.

Pith tools