Pith. sign in

REVIEW 8 cited by

MegaMath: Pushing the Limits of Open Math Corpora

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02807 v1 pith:RDUGTKHL submitted 2025-04-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords datacodeopenmath-relatedmegamathcorpuslargemath
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mathematical reasoning is a cornerstone of human intelligence and a key benchmark for advanced capabilities in large language models (LLMs). However, the research community still lacks an open, large-scale, high-quality corpus tailored to the demands of math-centric LLM pre-training. We present MegaMath, an open dataset curated from diverse, math-focused sources through following practices: (1) Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on the Internet. (2) Recalling Math-related code data: We identified high quality math-related code from large code training corpus, Stack-V2, further enhancing data diversity. (3) Exploring Synthetic data: We synthesized QA-style text, math-related code, and interleaved text-code blocks from web data or code data. By integrating these strategies and validating their effectiveness through extensive ablations, MegaMath delivers 371B tokens with the largest quantity and top quality among existing open math pre-training datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Latent Reasoning via Looped Language Models

    cs.CL 2025-10 unverdicted novelty 7.0 of 10

    Looped language models with latent iterative computation and entropy-regularized depth allocation achieve performance matching up to 12B standard LLMs through superior knowledge manipulation.

  2. DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A learned orchestrator builds per-example drop/untouch/clean pipelines over noise pruning and instruction-conditioned rewriting, improving from-scratch and math continued pretraining over fixed curation methods.

  3. OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A two-stage mid-training recipe on math-heavy corpora turns Llama-3.2 base models into ones whose RL math performance matches Qwen2.5 at the same size.

  4. Essential-Web v1.0: 24T tokens of organized web data

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 24T-token web corpus with automatic document-level taxonomy labels enables competitive domain-specific datasets via simple filters.

  5. Large-Scale Diverse Synthesis for Mid-Training

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    BoostQA is a 100B-token synthesized QA corpus whose mid-training on a 40B-token subset improves Llama-3 8B by 12.74% on average across MMLU and CMMLU and reaches top average performance on 12 benchmarks.

  6. LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    A knowledge point graph walk synthesizes a 50B token QA dataset that reportedly lifts Llama-3 8B average MMLU and CMMLU scores by 11.51%.

  7. SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Two new MoE language models and a co-designed inference engine claim 20+ tokens/s CPU decoding under 1-8 GB memory with benchmark scores comparable to much larger models.

  8. JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models

    cs.CL 2025-07 conditional novelty 4.0 of 10

    JT-Math-8B, an open 8B model family trained with a multi-stage math-focused pipeline, reports math benchmark averages above o1-mini and several 7B open models.

Pith tools