REVIEW 2 cited by
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
High-quality, large-scale corpora are the cornerstone of building foundation models. In this work, we introduce MathPile, a diverse and high-quality math-centric corpus comprising about 9.5 billion tokens. Throughout its creation, we adhered to the principle of "less is more", firmly believing in the supremacy of data quality over quantity, even in the pre-training phase. Our meticulous data collection and processing efforts included a complex suite of preprocessing, prefiltering, language identification, cleaning, filtering, and deduplication, ensuring the high quality of our corpus. Furthermore, we performed data contamination detection on downstream benchmark test sets to eliminate duplicates and conducted continual pre-training experiments, booting the performance on common mathematical reasoning benchmarks. We aim for our MathPile to boost language models' mathematical reasoning abilities and open-source its different versions and processing scripts to advance the field.
Forward citations
Cited by 2 Pith papers
-
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
StepHint improves RLVR math reasoning by giving the model multiple prefix-level hints drawn from correct chains generated by stronger models, beating several RLVR baselines on six math benchmarks and two out-of-domain sets.
-
MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning
MDPO applies a SimPO-style length-normalized reward to preference pairs built at solution, inference, and step granularities, yielding small accuracy gains on math benchmarks.
Discussion (0). Continue with ORCID to comment.