Pith. sign in

REVIEW 7 cited by

Augmenting Math Word Problems via Iterative Question Composing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.09003 v5 pith:6YNY7YXY submitted 2024-01-17 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords composingmathmmiqcmodelsproblemsdataimprovementiterative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the advancements in large language models (LLMs) for mathematical reasoning, solving competition-level math problems remains a significant challenge, especially for open-source LLMs without external tools. We introduce the MMIQC dataset, comprising a mixture of processed web data and synthetic question-response pairs, aimed at enhancing the mathematical reasoning capabilities of base language models. Models fine-tuned on MMIQC consistently surpass their counterparts in performance on the MATH benchmark across various model sizes. Notably, Qwen-72B-MMIQC achieves a 45.0% accuracy, exceeding the previous open-source state-of-the-art by 8.2% and outperforming the initial version GPT-4 released in 2023. Extensive evaluation results on Hungarian high school finals suggest that such improvement can generalize to unseen data. Our ablation study on MMIQC reveals that a large part of the improvement can be attributed to our novel augmentation method, Iterative Question Composing (IQC), which involves iteratively composing new questions from seed problems using an LLM and applying rejection sampling through another LLM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Turn On-Policy Distillation with Prefix Replay

    cs.LG 2026-07 conditional novelty 6.0 of 10

    ReOPD offline-distills multi-turn agentic LLMs via teacher-prefix replay plus step-decay sampling, matching online OPD accuracy at ≥4× speed with zero tool calls.

  2. SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Filtering self-generated math problems by a model's own solve-rate improves that model's MATH accuracy from 38% to 47% and helps out-of-distribution generalization when data is diverse.

  3. Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

    cs.LG 2026-08 conditional novelty 5.0 of 10

    A self-evolving curriculum that retrains a language model on variants of problems it can mostly get right lifts AIME pass@1 from 5.6% to 16.5%, beating static augmentation under the same data budget.

  4. A Survey of LLM $\times$ DATA

    cs.DB 2025-05 conditional novelty 5.0 of 10

    A comprehensive survey of the bidirectional links between LLMs and data management, organized as DATA4LLM and LLM4DATA with a new 'IaaS' data-quality framework.

  5. Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A process reward model trained with Error Propagation and Error Cessation labels from an o1 judge outperforms existing PRMs on math selection and step-level scoring.

  6. Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

    cs.CL 2025-02 reject novelty 5.0 of 10

    OREAL shows that outcome-reward RL with best-of-N positive behavior cloning, negative reward shaping, and token-level reweighting reaches state-of-the-art MATH-500 accuracy at 7B and 32B scale.

  7. A Survey on Large Language Models for Mathematical Reasoning

    cs.AI 2025-06 conditional novelty 1.0 of 10

    Recent advances in LLM mathematical reasoning are organized into comprehension and generation phases, covering methods from prompting to test-time scaling.

Pith tools