REVIEW 8 cited by
MegaMath: Pushing the Limits of Open Math Corpora
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Mathematical reasoning is a cornerstone of human intelligence and a key benchmark for advanced capabilities in large language models (LLMs). However, the research community still lacks an open, large-scale, high-quality corpus tailored to the demands of math-centric LLM pre-training. We present MegaMath, an open dataset curated from diverse, math-focused sources through following practices: (1) Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on the Internet. (2) Recalling Math-related code data: We identified high quality math-related code from large code training corpus, Stack-V2, further enhancing data diversity. (3) Exploring Synthetic data: We synthesized QA-style text, math-related code, and interleaved text-code blocks from web data or code data. By integrating these strategies and validating their effectiveness through extensive ablations, MegaMath delivers 371B tokens with the largest quantity and top quality among existing open math pre-training datasets.
Forward citations
Cited by 8 Pith papers
-
Scaling Latent Reasoning via Looped Language Models
Looped language models with latent iterative computation and entropy-regularized depth allocation achieve performance matching up to 12B standard LLMs through superior knowledge manipulation.
-
DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data
A learned orchestrator builds per-example drop/untouch/clean pipelines over noise pruning and instruction-conditioned rewriting, improving from-scratch and math continued pretraining over fixed curation methods.
-
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
A two-stage mid-training recipe on math-heavy corpora turns Llama-3.2 base models into ones whose RL math performance matches Qwen2.5 at the same size.
-
Essential-Web v1.0: 24T tokens of organized web data
A 24T-token web corpus with automatic document-level taxonomy labels enables competitive domain-specific datasets via simple filters.
-
Large-Scale Diverse Synthesis for Mid-Training
BoostQA is a 100B-token synthesized QA corpus whose mid-training on a 40B-token subset improves Llama-3 8B by 12.74% on average across MMLU and CMMLU and reaches top average performance on 12 benchmarks.
-
LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
A knowledge point graph walk synthesizes a 50B token QA dataset that reportedly lifts Llama-3 8B average MMLU and CMMLU scores by 11.51%.
-
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
Two new MoE language models and a co-designed inference engine claim 20+ tokens/s CPU decoding under 1-8 GB memory with benchmark scores comparable to much larger models.
-
JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models
JT-Math-8B, an open 8B model family trained with a multi-stage math-focused pipeline, reports math benchmark averages above o1-mini and several 7B open models.
Discussion (0). Sign in to comment.