Pith. sign in

REVIEW 18 cited by

Best Practices and Lessons Learned on Synthetic Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.07503 v2 pith:5SRO4UDH submitted 2024-04-11 cs.CL

classification cs.CL
keywords datasyntheticmodelsapplicationsartificialavailabilitybestbuild
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The success of AI models relies on the availability of large, diverse, and high-quality datasets, which can be challenging to obtain due to data scarcity, privacy concerns, and high costs. Synthetic data has emerged as a promising solution by generating artificial data that mimics real-world patterns. This paper provides an overview of synthetic data research, discussing its applications, challenges, and future directions. We present empirical evidence from prior art to demonstrate its effectiveness and highlight the importance of ensuring its factuality, fidelity, and unbiasedness. We emphasize the need for responsible use of synthetic data to build more powerful, inclusive, and trustworthy language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A learned orchestrator builds per-example drop/untouch/clean pipelines over noise pruning and instruction-conditioned rewriting, improving from-scratch and math continued pretraining over fixed curation methods.

  2. Search Hardness-Aware LLM-Based Problem Formulation for Expensive Simulation-Driven Design

    cs.NE 2026-07 conditional novelty 6.0 of 10

    SHA-PF uses initial simulation data to select a 'hard but promising' anchor satisfaction state and evolves LLM-generated formulations that prioritize it, reaching target designs with fewer expensive simulations on ant...

  3. Synthetic Interaction Data for Scalable Personalization in Large Language Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    PersonaGym simulates noisy multi-turn user–assistant interactions to build PersonaAtlas, and PPOpt learns to rewrite user prompts from interaction history, improving judged personalization on synthetic benchmarks.

  4. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A taxonomy-guided retrieval-augmented framework generates CultureSynth-7, a multilingual cultural QA benchmark, and its evaluation of 14 LLMs suggests cultural competence emerges around 3B parameters.

  5. Synthetic CVs To Build and Test Fairness-Aware Hiring Tools

    cs.CY 2025-08 conditional novelty 6.0 of 10

    A new synthetic CV dataset, generated from donated real CVs, is proposed as a benchmark for fairness-aware algorithmic hiring research.

  6. MDBench: A Synthetic Multi-Document Reasoning Benchmark Generated with Knowledge Guidance

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MDBench is a synthetically generated, knowledge-guided benchmark for multi-document QA on which frontier LLMs achieve only about 60% exact match.

  7. Seed-Coder: Let the Code Model Curate Data for Itself

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Seed-Coder shows that an LLM-trained quality scorer can filter 6T tokens of code data and yield 8B models that outperform similar-size open code models.

  8. EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    EasyDistill packages established LLM knowledge-distillation techniques into a single modular toolkit with released distilled models, datasets, and Alibaba Cloud integration.

  9. GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A knowledge-graph-guided method that scores an LLM's knowledge gaps and generates atomic, aggregated, and multi-hop QA pairs, improving closed-book QA after fine-tuning.

  10. OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs

    cs.AI 2025-05 conditional novelty 6.0 of 10

    OrgAccess, a 70k-query synthetic RBAC benchmark, shows current LLMs including GPT-4.1 (F1 0.27 on the hardest split) struggle badly with permission adherence.

  11. Enterprise Large Language Model Evaluation Benchmark

    cs.AI 2025-06 reject novelty 5.0 of 10

    A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...

  12. Using Sign Language Production as Data Augmentation to enhance Sign Language Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Adding synthetic sign-language data produced by stitching, a GAN, or Gaussian splatting to the training set improves sign-language translation, with the largest gains for skeleton-pose models.

  13. Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training

    cs.AI 2025-05 conditional novelty 5.0 of 10

    Reinforcing the two experts most correlated with thinking tokens improves reasoning accuracy and efficiency in MoE large reasoning models, with gains of up to 10 points on AIME benchmarks.

  14. Generative AI at the Crossroads: Light Bulb, Dynamo, or Microscope?

    econ.GN 2025-05 accept novelty 5.0 of 10

    GenAI is likely both a general-purpose technology and an invention of a method of invention, supporting a future rise in productivity levels with uncertain growth effects.

  15. Salamandra Technical Report

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.

  16. A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations

    cs.CL 2025-07 conditional novelty 4.0 of 10

    PATTR adds a target-length penalty to the Type-Token Ratio, producing a lexical diversity score with tunable, reduced short-text bias for LLM synthetic data.

  17. Does Prompt Design Impact Quality of Data Imputation by LLMs?

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Group-wise CSV prompts with correlation-based column pruning reduce LLM imputation prompt size while roughly maintaining or slightly improving classifier-based imputation quality on two imbalanced datasets.

  18. Unlocking the Potential of Large Language Models in the Nuclear Industry with Synthetic Data

    cs.CL 2025-06 conditional novelty 2.0 of 10

    A pipeline converts CANDU textbook chapters into synthetic QA pairs using LLMs, embedding clustering, and similarity metrics, with no downstream validation.

Pith tools