Pith. sign in

REVIEW 3 cited by

Using GANs for Sharing Networked Time Series Data: Challenges, Initial Promise, and Open Questions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.13403 v5 pith:REC7QLV6 submitted 2019-09-30 cs.LG cs.DCcs.NIstat.ML

classification cs.LGcs.DCcs.NIstat.ML
keywords challengesdatafidelityprivacysharingdatasetsganswork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Limited data access is a longstanding barrier to data-driven research and development in the networked systems community. In this work, we explore if and how generative adversarial networks (GANs) can be used to incentivize data sharing by enabling a generic framework for sharing synthetic datasets with minimal expert knowledge. As a specific target, our focus in this paper is on time series datasets with metadata (e.g., packet loss rate measurements with corresponding ISPs). We identify key challenges of existing GAN approaches for such workloads with respect to fidelity (e.g., long-term dependencies, complex multidimensional relationships, mode collapse) and privacy (i.e., existing guarantees are poorly understood and can sacrifice fidelity). To improve fidelity, we design a custom workflow called DoppelGANger (DG) and demonstrate that across diverse real-world datasets (e.g., bandwidth measurements, cluster requests, web sessions) and use cases (e.g., structural characterization, predictive modeling, algorithm comparison), DG achieves up to 43% better fidelity than baseline models. Although we do not resolve the privacy problem in this work, we identify fundamental challenges with both classical notions of privacy and recent advances to improve the privacy properties of GANs, and suggest a potential roadmap for addressing these challenges. By shedding light on the promise and challenges, we hope our work can rekindle the conversation on workflows for data sharing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthetic Time Series Generation via Complex Networks

    cs.LG 2026-01 conditional novelty 4.0 of 10

    A first-order Markov chain on empirical quantile bins — built from the original series and sampled to produce new series — reproduces marginal distributions and lag-1 correlations but cannot capture long-range or high...

  2. Federated Diffusion Modeling with Differential Privacy for Tabular Data Synthesis

    cs.LG 2024-12 reject novelty 4.0 of 10

    DP-FedTabDiff wraps an existing federated tabular diffusion model with per-client DP-SGD and reports how the privacy budget, number of clients, and local update count affect synthetic data quality and empirical privacy risk.

  3. Synthetic Time Series Data Generation for Healthcare Applications: A PCG Case Study

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A comparison of three generative models for synthetic PCG signals reports low distribution-distance scores, but the evaluation lacks held-out splits, baselines, and error bars.

Pith tools