REVIEW 8 cited by
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have shown excellent mastering of human language, but still struggle in real-world applications that require mathematical problem-solving. While many strategies and datasets to enhance LLMs' mathematics are developed, it remains a challenge to simultaneously maintain and improve both language and mathematical capabilities in deployed LLM systems.In this work, we tailor the Self-Critique pipeline, which addresses the challenge in the feedback learning stage of LLM alignment. We first train a general Math-Critique model from the LLM itself to provide feedback signals. Then, we sequentially employ rejective fine-tuning and direct preference optimization over the LLM's own generations for data collection. Based on ChatGLM3-32B, we conduct a series of experiments on both academic and our newly created challenging dataset, MathUserEval. Results show that our pipeline significantly enhances the LLM's mathematical problem-solving while still improving its language ability, outperforming LLMs that could be two times larger. Related techniques have been deployed to ChatGLM\footnote{\url{https://chatglm.cn}}, an online serving LLM. Related evaluation dataset and scripts are released at \url{https://github.com/THUDM/ChatGLM-Math}.
Forward citations
Cited by 8 Pith papers
-
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
Step-aligned in-context learning with a first-try retrieval strategy improves LLM mathematical reasoning over problem-level few-shot prompting on multiple benchmarks.
-
DIVE: Diversified Iterative Self-Improvement
DIVE combines global sample pooling with diversity-aware data selection to counter output diversity collapse in iterative preference learning for LLMs.
-
System-2 Mathematical Reasoning via Enriched Instruction Tuning
Enriched Instruction Tuning (EIT) uses GPT-4 to add planning and missing reasoning steps to human-annotated math solutions, and fine-tuning LLaMA-2 on this data yields 84.1% on GSM8K and 32.5% on MATH.
-
Teaching LLMs to Refine with Tools
CaP trains LLMs to fix chain-of-thought math solutions by producing program-of-thought code, and shows that DPO preference optimization is essential for the refinement to actually improve accuracy.
-
Mars-PO: Multi-Agent Reasoning System Preference Optimization
A multi-agent preference optimization method that uses pooled correct answers from several LLMs as shared positives and each model's own errors as negatives improves math reasoning accuracy on GSM8K and MATH.
-
Diving into Self-Evolving Training for Multimodal Reasoning
M-STAR, a self-evolving training recipe combining continuous updates, a process-reward-model reranker, and adaptive sampling temperature, improves multimodal reasoning on several benchmarks across three vision-languag...
-
ICH-Qwen: A Large Language Model Towards Chinese Intangible Cultural Heritage
They fine-tuned Qwen2.5-7B on Chinese intangible cultural heritage texts to build ICH-Qwen, and report n-gram metric wins over general LLMs on 100-sample ICH QA tasks.
-
End-to-End Bangla AI for Solving Math Olympiad Problem Benchmark: Leveraging Large Language Model Using Integrated Approach
Fine-tuning Qwen2.5-7B with translated math datasets plus retrieval and tool-integrated reasoning yields 71/100 on a Bangla math olympiad test set, versus 77/100 for a larger base model.
Discussion (0). Continue with ORCID to comment.