Pith. sign in

REVIEW 6 cited by

Comprehensive Exploration of Synthetic Data Generation: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.02524 v2 pith:5I5HETH6 submitted 2024-01-04 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords datamodelsgenerationmodelsyntheticcommoncomprehensiveexploration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent years have witnessed a surge in the popularity of Machine Learning (ML), applied across diverse domains. However, progress is impeded by the scarcity of training data due to expensive acquisition and privacy legislation. Synthetic data emerges as a solution, but the abundance of released models and limited overview literature pose challenges for decision-making. This work surveys 417 Synthetic Data Generation (SDG) models over the last decade, providing a comprehensive overview of model types, functionality, and improvements. Common attributes are identified, leading to a classification and trend analysis. The findings reveal increased model performance and complexity, with neural network-based approaches prevailing, except for privacy-preserving data generation. Computer vision dominates, with GANs as primary generative models, while diffusion models, transformers, and RNNs compete. Implications from our performance evaluation highlight the scarcity of common metrics and datasets, making comparisons challenging. Additionally, the neglect of training and computational costs in literature necessitates attention in future research. This work serves as a guide for SDG model selection and identifies crucial areas for future exploration.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 17 citations worldwide. Full citation record

  1. RaMark: Radioactive Watermarking for Generated Tabular Data

    cs.CR 2026-07 conditional novelty 7.0 of 10

    A sinusoidal dependency embedded as part of the tabular distribution remains detectable after generative retraining and data-modification attacks while utility is preserved.

  2. Reading a Ruler in the Wild

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RulerNet detects centimeter marks on rulers in natural images and estimates pixel-per-centimeter scale via a learned geometric progression model, outperforming prior OCR-based and frequency-based baselines.

  3. Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)

    cs.MA 2026-01 reject novelty 5.0 of 10

    AoD pairs a frozen diffusion language model with two LLM agents that iteratively rewrite prompts from natural-language feedback, reporting better JSON diversity and validity, though the claimed RL mechanism and theore...

  4. RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Fine-tuning a vision-language model on a synthetic 4-clue dataset and adding a self-critique inference loop improves robustness to domain shifts in object recognition.

  5. Enterprise Large Language Model Evaluation Benchmark

    cs.AI 2025-06 reject novelty 5.0 of 10

    A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...

  6. Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques

    cs.LG 2025-07 conditional novelty 3.0 of 10

    A survey that categorizes tabular data synthesis by generation objectives and adds a benchmark comparison of six models on Adult and CreditRisk.

Pith tools