Pith. sign in

REVIEW 1 cited by

Benchmarking the Fidelity and Utility of Synthetic Relational Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03411 v1 pith:RD7PKZ5Q submitted 2024-10-04 cs.DB cs.LG

classification cs.DBcs.LG
keywords databenchmarkingmethodsrelationalsynthesizingsyntheticutilityevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Synthesizing relational data has started to receive more attention from researchers, practitioners, and industry. The task is more difficult than synthesizing a single table due to the added complexity of relationships between tables. For the same reason, benchmarking methods for synthesizing relational data introduces new challenges. Our work is motivated by a lack of an empirical evaluation of state-of-the-art methods and by gaps in the understanding of how such an evaluation should be done. We review related work on relational data synthesis, common benchmarking datasets, and approaches to measuring the fidelity and utility of synthetic data. We combine the best practices and a novel robust detection approach into a benchmarking tool and use it to compare six methods, including two commercial tools. While some methods are better than others, no method is able to synthesize a dataset that is indistinguishable from original data. For utility, we typically observe moderate correlation between real and synthetic data for both model predictive performance and feature importance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TabQueryBench: A Query-Centric Benchmark for Synthetic Tabular Data

    cs.DB 2026-07 accept novelty 7.0 of 10

    Across 49 datasets and 11 generators, distance-based fidelity overstates synthetic tabular quality: best query-centric score is only 0.75, with systematic failures on high-cardinality support, local conditionals, and ...

Pith tools