REVIEW 13 cited by
On the Diversity of Synthetic Data and its Impact on Training Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rise of Large Language Models (LLMs) has accentuated the need for diverse, high-quality pre-training data. Synthetic data emerges as a viable solution to the challenges of data scarcity and inaccessibility. While previous literature has focused predominantly on the quality and quantity of real data, our work enables the measurement of diversity in synthetic data and explores its impact on LLM performance. We study the downstream effects of synthetic data diversity during both the pre-training and fine-tuning stages by introducing a new diversity metric, \textit{LLM cluster-agent}, designed to evaluate the diversity of synthetic datasets. Through a series of controlled experiments with models of 350M and 1.4B parameters, we demonstrate that the proposed cluster-based LLM scoring of diversity correlates positively with both pre-training and supervised fine-tuning performance. Our findings also reveal that synthetic data diversity in pre-training affects supervised fine-tuning more significantly than pre-training itself, even for smaller models. We hope this study advances our understanding of the optimal use of synthetic data in LLM training and opens new avenues for efficient data generation processes.
Forward citations
Cited by 13 Pith papers
-
Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation
Gemma 3 27B and Aya Expanse 32B are the strongest multilingual synthetic-data teachers; model scale does not predict effectiveness while prompt diversity, length and response fluency do.
-
Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).
-
Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
Verifier-filtered synthetic retraining improves linear-regression estimates in the short term but converges to the verifier's knowledge center, so sustained improvement requires an unbiased verifier.
-
One Joke to Rule them All? On the (Im)possibility of Generalizing Humor
LLMs fine-tuned on one to three humor datasets transfer partially to unseen humor types (up to 75% accuracy); diverse training helps modestly, and dad jokes enable transfer best but resist it as a target.
-
Towards Integrated Alignment
LLM-generated textbook-style forget sets, produced from a domain name alone, achieve unlearning performance comparable to expert-curated datasets in biosecurity, cybersecurity, and Harry Potter benchmarks.
-
Decoding Machine Translationese in English-Chinese News: LLMs vs. NMTs
Machine-translated English-to-Chinese news differs from original Chinese news in measurable ways, and LLM and NMT outputs can be partially but not fully distinguished by linguistic features.
-
What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning
Moderately diverse LLM-generated data can improve fine-tuned model performance in low-data settings when distribution shift is minimal, while high diversity or large distribution shift hurts.
-
Improving Multilingual Math Reasoning for African Languages
SFT on translated OpenMathInstruct data outperforms directly generated synthetic data for math in African languages, and combining both yields the best AfriMGSM scores.
-
From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning
Controlling the variety of mid-frequency response tokens during SFT dataset construction correlates more strongly with downstream model performance than instruction-level diversity strategies.
-
Diversity and Inclusion in AI: Insights from a Survey of AI/ML Practitioners
A survey of 61 AI/ML practitioners finds that while most believe diverse teams and data reduce bias, actual practices like bias audits and post-development D&I checks are inconsistent and often missing.
-
A Survey of LLM $\times$ DATA
A comprehensive survey of the bidirectional links between LLMs and data management, organized as DATA4LLM and LLM4DATA with a new 'IaaS' data-quality framework.
-
Facts Do Care About Your Language: Assessing Answer Quality of Multilingual LLMs
A small evaluation of Llama 3.1 shows factuality in school-level question answering degrades with decreasing language speaker count, though the statistical support is weakened by methodological issues.
-
LLM Harms: A Taxonomy and Discussion
This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.
Discussion (0). Sign in to comment.