Pith. sign in

REVIEW 5 cited by

How to Synthesize Text Data without Model Collapse?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.14689 v3 pith:VMFX7YHV submitted 2024-12-19 cs.CL cs.AIcs.LG

How to Synthesize Text Data without Model Collapse?

classification cs.CL cs.AIcs.LG
keywords datamodelsyntheticcollapseeditingmodelsperformanceconduct
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-$\{n\}$ models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves model performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs

    cs.SE 2026-06 unverdicted novelty 6.0

    Experiments across code LLMs show no-review collapses fastest, human-gated filters slow collapse, and AI self-gates lose effect over time, degenerating to ungated self-training under self-confirming acceptance as prov...

  2. A Study of LLMs' Preferences for Libraries and Programming Languages

    cs.SE 2025-03 unverdicted novelty 6.0

    Empirical study of eight LLMs finds overuse of popular libraries like NumPy in up to 45% of unnecessary cases and strong default preference for Python even when suboptimal.

  3. Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks

    cs.LG 2025-02 unverdicted novelty 6.0

    Empirical study across 10 tasks showing bias inheritance from LLM-augmented data harms related downstream performance, with three misalignment factors and three mitigation strategies identified.

  4. Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities

    cs.LG 2026-05 conditional novelty 4.0

    Model collapse threatens AI democratization by disproportionately degrading data and efficiency for low-resource communities.

  5. Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities

    cs.LG 2026-05 unverdicted novelty 3.0

    Model collapse threatens AI democratization by disproportionately impacting low-resource and marginalized communities through reduced training efficiency and data distributions skewed away from distribution tails.