Pith. sign in

REVIEW 8 cited by

ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.02893 v1 pith:6BZJIS5C submitted 2024-04-03 cs.CL

classification cs.CL
keywords languagellmsmathematicalpipelineproblem-solvingchallengechatglmchatglm-math
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have shown excellent mastering of human language, but still struggle in real-world applications that require mathematical problem-solving. While many strategies and datasets to enhance LLMs' mathematics are developed, it remains a challenge to simultaneously maintain and improve both language and mathematical capabilities in deployed LLM systems.In this work, we tailor the Self-Critique pipeline, which addresses the challenge in the feedback learning stage of LLM alignment. We first train a general Math-Critique model from the LLM itself to provide feedback signals. Then, we sequentially employ rejective fine-tuning and direct preference optimization over the LLM's own generations for data collection. Based on ChatGLM3-32B, we conduct a series of experiments on both academic and our newly created challenging dataset, MathUserEval. Results show that our pipeline significantly enhances the LLM's mathematical problem-solving while still improving its language ability, outperforming LLMs that could be two times larger. Related techniques have been deployed to ChatGLM\footnote{\url{https://chatglm.cn}}, an online serving LLM. Related evaluation dataset and scripts are released at \url{https://github.com/THUDM/ChatGLM-Math}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning

    cs.CL 2025-01 conditional novelty 7.0 of 10

    Step-aligned in-context learning with a first-try retrieval strategy improves LLM mathematical reasoning over problem-level few-shot prompting on multiple benchmarks.

  2. DIVE: Diversified Iterative Self-Improvement

    cs.CL 2025-01 conditional novelty 6.0 of 10

    DIVE combines global sample pooling with diversity-aware data selection to counter output diversity collapse in iterative preference learning for LLMs.

  3. System-2 Mathematical Reasoning via Enriched Instruction Tuning

    cs.AI 2024-12 conditional novelty 6.0 of 10

    Enriched Instruction Tuning (EIT) uses GPT-4 to add planning and missing reasoning steps to human-annotated math solutions, and fine-tuning LLaMA-2 on this data yields 84.1% on GSM8K and 32.5% on MATH.

  4. Teaching LLMs to Refine with Tools

    cs.CL 2024-12 conditional novelty 6.0 of 10

    CaP trains LLMs to fix chain-of-thought math solutions by producing program-of-thought code, and shows that DPO preference optimization is essential for the refinement to actually improve accuracy.

  5. Mars-PO: Multi-Agent Reasoning System Preference Optimization

    cs.AI 2024-11 conditional novelty 6.0 of 10

    A multi-agent preference optimization method that uses pooled correct answers from several LLMs as shared positives and each model's own errors as negatives improves math reasoning accuracy on GSM8K and MATH.

  6. Diving into Self-Evolving Training for Multimodal Reasoning

    cs.CL 2024-12 conditional novelty 5.0 of 10

    M-STAR, a self-evolving training recipe combining continuous updates, a process-reward-model reranker, and adaptive sampling temperature, improves multimodal reasoning on several benchmarks across three vision-languag...

  7. ICH-Qwen: A Large Language Model Towards Chinese Intangible Cultural Heritage

    cs.CL 2025-05 reject novelty 4.0 of 10

    They fine-tuned Qwen2.5-7B on Chinese intangible cultural heritage texts to build ICH-Qwen, and report n-gram metric wins over general LLMs on 100-sample ICH QA tasks.

  8. End-to-End Bangla AI for Solving Math Olympiad Problem Benchmark: Leveraging Large Language Model Using Integrated Approach

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Fine-tuning Qwen2.5-7B with translated math datasets plus retrieval and tool-integrated reasoning yields 71/100 on a Bangla math olympiad test set, versus 77/100 for a larger base model.

Pith tools