Pith. sign in

REVIEW 3 cited by

RedStone: Curating General, Code, Math, and QA Data for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.03398 v1 pith:4HWMKNXZ submitted 2024-12-04 cs.CL

classification cs.CL
keywords redstonedatadatasetspre-trainingcommoncrawllanguagellms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-training Large Language Models (LLMs) on high-quality, meticulously curated datasets is widely recognized as critical for enhancing their performance and generalization capabilities. This study explores the untapped potential of Common Crawl as a comprehensive and flexible resource for pre-training LLMs, addressing both general-purpose language understanding and specialized domain knowledge. We introduce RedStone, an innovative and scalable pipeline engineered to extract and process data from Common Crawl, facilitating the creation of extensive and varied pre-training datasets. Unlike traditional datasets, which often require expensive curation and domain-specific expertise, RedStone leverages the breadth of Common Crawl to deliver datasets tailored to a wide array of domains. In this work, we exemplify its capability by constructing pre-training datasets across multiple fields, including general language understanding, code, mathematics, and question-answering tasks. The flexibility of RedStone allows for easy adaptation to other specialized domains, significantly lowering the barrier to creating valuable domain-specific datasets. Our findings demonstrate that Common Crawl, when harnessed through effective pipelines like RedStone, can serve as a rich, renewable source of pre-training data, unlocking new avenues for domain adaptation and knowledge discovery in LLMs. This work also underscores the importance of innovative data acquisition strategies and highlights the role of web-scale data as a powerful resource in the continued evolution of LLMs. RedStone code and data samples will be publicly available at \url{https://aka.ms/redstone}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction

    cs.LG 2026-01 conditional novelty 5.0 of 10

    Distill-then-Replace builds task-specific hybrid attention LLMs by distilling each full-attention block into a linear counterpart and greedily replacing layers under a validation-performance constraint.

  2. Large-Scale Diverse Synthesis for Mid-Training

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    BoostQA is a 100B-token synthesized QA corpus whose mid-training on a 40B-token subset improves Llama-3 8B by 12.74% on average across MMLU and CMMLU and reaches top average performance on 12 benchmarks.

  3. Data Efficacy for Language Model Training

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Ordering training data by a gradient-based score, using a folding scheme that interleaves multiple curriculum passes, improves small-scale LM accuracy by roughly 1.5 to 2 points on average benchmarks.

Pith tools