Pith. sign in

REVIEW 4 cited by

Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.20796 v1 pith:HE742ZBL submitted 2024-10-28 cs.CL

classification cs.CL
keywords datapre-trainingrephrasingmodelresultsdifferentperformancepipeline
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently published work on rephrasing natural text data for pre-training LLMs has shown promising results when combining the original dataset with the synthetically rephrased data. We build upon previous work by replicating existing results on C4 and extending them with our optimized rephrasing pipeline to the English, German, Italian, and Spanish Oscar subsets of CulturaX. Our pipeline leads to increased performance on standard evaluation benchmarks in both the mono- and multilingual setup. In addition, we provide a detailed study of our pipeline, investigating the choice of the base dataset and LLM for the rephrasing, as well as the relationship between the model size and the performance after pre-training. By exploring data with different perceived quality levels, we show that gains decrease with higher quality. Furthermore, we find the difference in performance between model families to be bigger than between different model sizes. This highlights the necessity for detailed tests before choosing an LLM to rephrase large amounts of data. Moreover, we investigate the effect of pre-training with synthetic data on supervised fine-tuning. Here, we find increasing but inconclusive results that highly depend on the used benchmark. These results (again) highlight the need for better benchmarking setups. In summary, we show that rephrasing multilingual and low-quality data is a very promising direction to extend LLM pre-training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Memorization vs. Reasoning: Updating LLMs with New Knowledge

    cs.CL 2025-04 conditional novelty 7.0 of 10

    KUP and MCT: a new benchmark and training method showing LLMs can memorize post-cutoff knowledge updates but fail to reason over them in indirect tests.

  2. Reformulation for Pretraining Data Augmentation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    MGA reformulates existing high-quality text into diverse genre-audience variants, producing a 770B-token corpus that improves LLM pretraining under data-constrained, high-repetition conditions.

  3. Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A 1.6B Arabic language model trained with synthetic multiple-choice instruction data outperforms 7B-13B models on several Arabic multiple-choice benchmarks.

  4. Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration

    cs.LG 2025-09 conditional novelty 4.0 of 10

    Constraining LLM rewriting with a biomedical NER model improves medical entity preservation and reduces hallucinations in synthetic clinical notes, with modest downstream gains on MIMIC-III tasks.

Pith tools