Pith. sign in

REVIEW 3 cited by

Recursive Learning Without Collapse: A Weighting-Based Stabilization Framework

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.18049 v6 pith:4USJJQBM submitted 2025-02-25 stat.ML cs.LG

classification stat.MLcs.LG
keywords datamodelsynthetictrainingmodelsperformancerealgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent studies identified an intriguing phenomenon in recursive generative model training known as model collapse, where models trained on data generated by previous models exhibit severe performance degradation. Addressing this issue and developing more effective training strategies have become central challenges in generative model research. In this paper, we investigate this phenomenon within a novel framework, where generative models are iteratively trained on a combination of newly collected real data and synthetic data from the previous training step. To develop an optimal training strategy for integrating real and synthetic data, we evaluate the performance of a weighted training scheme in various scenarios, including Gaussian distribution estimation, generalized linear models, and nonparametric estimation. We theoretically characterize the impact of the mixing proportion and weighting scheme of synthetic data on the final model's performance. Our key finding is that, across different settings, the optimal weighting scheme under different proportions of synthetic data asymptotically follows a unified expression, revealing a fundamental trade-off between leveraging synthetic data and model performance. In some cases, the optimal weight assigned to real data corresponds to the reciprocal of the golden ratio. Finally, we validate our theoretical results on extensive simulated datasets and a real tabular dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence

    stat.ML 2025-10 conditional novelty 6.0 of 10

    Verifier-filtered synthetic retraining improves linear-regression estimates in the short term but converges to the verifier's knowledge center, so sustained improvement requires an unbiased verifier.

  2. A Probabilistic Perspective on Model Collapse

    stat.ML 2025-05 conditional novelty 6.0 of 10

    Recursive training on synthetic data avoids model collapse if the per-generation sample size grows superlinearly, and the probability that synthetic retraining improves over real-data training is always below one half.

  3. LLM Web Dynamics: Tracing Model Collapse in a Network of LLMs

    cs.LG 2025-05 conditional novelty 4.0 of 10

    Under a shared retrieval-augmented memory, multiple LLMs' outputs converge to near-identical semantic answers, and the analogous Gaussian mixture system is proven to collapse.

Pith tools