REVIEW 2 major objections 5 minor 5 cited by
A Survey on Large Language Models for Mathematical Reasoning
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This survey argues that mathematical reasoning in LLMs advances through two phases—comprehension, built by pretraining, and answer generation, led by chain-of-thought reasoning—and that this split organizes the field's methods and open…
desk verdict A useful, up-to-date survey of LLM math reasoning that is let down by a framework that doesn't actually organize the content and some fixable production errors. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is the two-phase division of comprehension and answer generation. Comprehension names the mathematical understanding internalized during pretraining; answer generation names the production of solutions, from direct prediction to structured step-by-step reasoning. Within the generation phase, CoT prompting is the key mechanism: intermediate tokens give transformers greater computational reach, so CoT is the load-bearing object connecting the two phases.
What would settle it
A controlled comparison of models matched on pretraining but differing only in CoT training, and vice versa, would show whether one phase alone drives the accuracy gains on benchmarks like MATH or AIME.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that LLMs acquire mathematical competence through comprehension first and answer generation second. Comprehension comes from pretraining on large, curated mathematical corpora; answer generation is the reasoning behavior elicited at inference or fine-tuning time, whose keystone is CoT prompting. The survey reads recent advances as a sequence of ways to strengthen the generation phase, and it gathers evidence that these methods largely unlock latent capabilities rather than create new ones.
Load-bearing premise
The survey's map rests on the assumption that LLM mathematical ability cleanly splits into a comprehension phase and an answer-generation phase.
Editorial extensions
If this is right
- CoT is more than a prompting trick; it lets models emulate structured procedures, so extending CoT with long reasoning or structural search should keep improving accuracy on multi-step mathematics problems.
- Reinforcement learning with rule-based rewards improves Pass@1 and, with long CoT, Pass@k, but mostly by activating capabilities already present in the base model.
- Test-time scaling offers a practical alternative to scaling model parameters for mathematics benchmarks.
Reading between the lines
- If reinforcement learning mainly unlocks latent capability, then measuring a base model's Pass@k before training could predict how much RL can help, which the survey does not explicitly propose.
- The comprehension-versus-generation framing could transfer to non-mathematical reasoning domains, predicting that CoT and search help most when the domain's knowledge is already in pretraining.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews research on mathematical reasoning in large language models (LLMs). It organizes the field around two proposed cognitive phases: comprehension (acquiring mathematical understanding through pretraining on diverse corpora) and answer generation (progressing from direct prediction to Chain-of-Thought reasoning). The paper covers background on pretraining, SFT, RL, and prompting; methods for boosting reasoning (SFT, RL, DPO, test-time structural search, self-improvement, external knowledge/tools); discussion of limitations; and future research directions. It also provides tables summarizing pretraining datasets, SFT datasets, and evaluation benchmarks.
Significance. The paper's compilation of recent work on long CoT, rule-based RL, test-time scaling, and self-improvement is timely, and its dataset/benchmark tables (Tables 1–3) are a useful reference. The discussion in Section 5 of RL as unlocking latent capabilities rather than creating new ones is a fair and well-supported observation. However, the central claim—that the paper consolidates fragmented developments into a coherent framework based on comprehension and generation—is not actually realized: the methods in Section 4 are never mapped onto the proposed phases. As a synthesis, the contribution is therefore currently incomplete, though the underlying survey content is largely accurate and potentially valuable after restructuring.
major comments (2)
- [Section 3 / Figure 2] The comprehension/generation dichotomy is asserted in the abstract and Section 3, and Figure 2 presents it as the organizing structure, but Section 4 never uses these categories. Section 4.1 (SFT), Section 4.2 (RL), Section 4.3 (test-time search), Section 4.4 (self-improvement), and Section 4.5 (external knowledge) are not classified as either comprehension or answer generation, and the text provides no explanation of how, say, RL with rule-based rewards or RAG relates to the two phases. Because the claimed 'coherent framework' (Section 1) is the paper's stated contribution, this omission is load-bearing: the framework is decorative rather than organizing.
- [Section 6] The roadmap at the start of Section 6 promises three directions but cites 'Sections 6 and 6', '(6)', and 'Section 6'—all self-references. The actual subsections are not numbered, so the reader cannot map the three promises (extending performance boundaries, enhancing efficiency, cross-domain generalization) onto the content that follows. This structural error undermines the paper's forward-looking contribution and needs to be fixed, either with proper subsection numbers and cross-references or by removing the broken pointers.
minor comments (5)
- [Section 4.1] The first sentence is grammatically incomplete: 'align a pretrained language model with high-quality, human-crafted by supervised learning' appears to be missing a noun such as 'data'.
- [Tables 1–3] The 'Synth.' column uses '%' to indicate non-synthetic datasets, which is confusing; use 'No' or a dash instead, since '%' reads as a percentage or a formatting artifact.
- [Figure 2] Figure 2 renders 'Comprehension and Generation (§3)' and 'Methods for Boosting Reasoning (§4)' as sibling boxes, visually contradicting the text's claim that Section 4 methods fit within the comprehension/generation framework.
- [References] Besta et al. 2024a and 2024b are the same paper (Graph of Thoughts, AAAI 2024) listed twice; similarly, Sprague et al. 2024a and 2024b are duplicates. These should be merged into single entries with consistent in-text citations.
- [Throughout] Model naming is inconsistent: 'DeepSeek R1' appears both as 'DeepSeek' and 'Deepseek', and 'Kimi K1.5' appears also as 'Kimi k1.5'. Please standardize.
Circularity Check
No significant circularity: the survey's two-phase framework and method review are organizational, with no derivation or prediction resting on its own inputs.
full rationale
This is a survey paper; it does not derive or predict results from fitted parameters or from a uniqueness theorem. Its central organizing claim—comprehension via pretraining/data and answer generation via CoT—is introduced as a high-level cognitive framing (Section 3) and used to structure the literature review; it is not obtained from the surveyed methods by construction. The self-citations in Section 6 (BWArea [Jia et al., 2024], CoLA [Jia et al., 2025]) and Section 4.4 (SIRLC [Pang et al., 2023]) are presented as promising research directions or as examples of self-improvement methods, but none of these citations carries the load of the survey's synthesis or of any factual claim about the field; removing them would not change the framework or the survey's conclusions. The potential weakness that Section 4 methods (SFT, RL, DPO, tree search, external tools) are not explicitly classified into the two phases is an organizational/coherence concern, not a circular one: no equation, definition, or fitted quantity reduces to its own input. The paper's own caveat in Section 5 that Pass@k may be an inadequate capability measure is a self-acknowledged limitation, not a circularity. No concrete circular step can be quoted, so the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The cited references accurately summarize the state of the art in LLM mathematical reasoning.
- ad hoc to paper The comprehension/generation dichotomy is a valid organizing principle for the field.
- domain assumption Benchmark scores cited (e.g., Grok 3 Beta on AIME 2024) are accurate and comparable.
Cite this review
Pith. "Pith review of A Survey on Large Language Models for Mathematical Reasoning." pith.science (2026). https://pith.science/paper/VUOVRYXS
@misc{pith2026250608446,
author = {Pith},
title = {Pith review of: A Survey on Large Language Models for Mathematical Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/VUOVRYXS}},
note = {Machine review of arXiv:2506.08446}
}
read the original abstract
Mathematical reasoning has long represented one of the most fundamental and challenging frontiers in artificial intelligence research. In recent years, large language models (LLMs) have achieved significant advances in this area. This survey examines the development of mathematical reasoning abilities in LLMs through two high-level cognitive phases: comprehension, where models gain mathematical understanding via diverse pretraining strategies, and answer generation, which has progressed from direct prediction to step-by-step Chain-of-Thought (CoT) reasoning. We review methods for enhancing mathematical reasoning, ranging from training-free prompting to fine-tuning approaches such as supervised fine-tuning and reinforcement learning, and discuss recent work on extended CoT and "test-time scaling". Despite notable progress, fundamental challenges remain in terms of capacity, efficiency, and generalization. To address these issues, we highlight promising research directions, including advanced pretraining and knowledge augmentation techniques, formal reasoning frameworks, and meta-generalization through principled learning paradigms. This survey tries to provide some insights for researchers interested in enhancing reasoning capabilities of LLMs and for those seeking to apply these techniques to other domains.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 5 Pith papers
-
FormalRx: Rectify and eXamine Semantic Failures in Autoformalization
FormalRx diagnoses Lean autoformalization failures with a 28-category SCI taxonomy and an 8B model that jointly predicts alignment, error type, location, and correction.
-
Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems
A fully automated pipeline produces proof-centric math benchmarks, demonstrated on algebraic geometry with 456 items, where leading LLMs score near 60 percent.
-
Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving
A multi-round, lemma-memory reasoning agent with hierarchical RL reaches reported gold-medal-level scores on Olympiad math benchmarks, though the proof-based scores are self-graded.
-
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
MMR-GRPO reweights GRPO rewards by semantic diversity, claiming ~48% fewer steps and ~70% less time while matching peak math-reasoning performance.
-
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
Post-training LLMs first on abstract number-free reasoning plans (CoMT), then with confidence-weighted rewards (CCRL), raises math accuracy by ~2-5 points over standard CoT-SFT+RL across four models.
Reference graph
Works this paper leans on
- [2]
-
[7]
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, J. Kernion, K. Ndousse, C. Olsson, D. Amodei, T. B. Brown, J. Clark, S. McCandlish, C. Olah, and J. Kaplan. A general language assistant as a laboratory for alignment. CoRR, abs/2112.00861,
-
[9]
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. Mc- Candlish, C. Olah, B. Mann, and J. Kaplan....
-
[11]
M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P . Nyczyk, and T. Hoefler. Graph of thoughts: Solving elaborate problems with large language models. In M. J. Wooldridge, J. G. Dy, and S. Natarajan, editors,Proceedings of the 38th AAAI Conference on Artificial Intelligence, 2024a. M. B...
-
[15]
Y. Chen, J. Benton, A. Radhakrishnan, J. Uesato, C. Denison, J. Schulman, A. Somani, P . Hase, M. Wagner, F. Roger, et al. Reasoning models don’t always say what they think.CoRR, abs/2505.05410, 2025c. 20 A Survey on Large Language Models for Mathematical Reasoning E. Chern, H. Zou, X. Li, J. Hu, K. Feng, J. Li, and P . Liu. Generative AI for math: Abel. ...
- [16]
-
[18]
D. Das, D. Banerjee, S. Aditya, and A. Kulkarni. MATHSENSEI: A tool-augmented large language model for mathematical reasoning.CoRR, abs/2402.17231,
-
[19]
A. Didolkar, A. Goyal, N. R. Ke, S. Guo, M. Valko, T. Lillicrap, D. Rezende, Y. Bengio, M. Mozer, and S. Arora. Metacognitive capabilities of LLMs: An exploration in mathematical problem solving.CoRR, abs/2405.12205,
Show all 119 references
-
[20]
Y. Ding, X. Shi, X. Liang, J. Li, Q. Zhu, and M. Zhang. Unleashing reasoning capability of LLMs via scalable question synthesis from scratch.CoRR, abs/2410.18693,
-
[21]
Dixit and T
P . Dixit and T. Oates. SBI-RAG: Enhancing math word problem solving for students through schema-based instruction and retrieval-augmented generation.CoRR, 2410.13293,
-
[22]
Dubey, A
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Rozière, B. Biron, B. Ta...
-
[24]
S. Feng, X. Kong, S. Ma, A. Zhang, D. Yin, C. Wang, R. Pang, and Y. Yang. Step-by-step reasoning for math problems via twisted sequential monte carlo.CoRR, abs/2410.01920, 2024b. C. R. Fletcher. Understanding and solving arithmetic word problems: A computer simulation.Behavior...
-
[27]
L. Gao, A. Madaan, S. Zhou, U. Alon, P . Liu, Y. Yang, J. Callan, and G. Neubig. PAL: Program-aided language models. InProceedings of the 40th International Conference on Machine Learning, 2023a. Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, H. Wang, and H. ...
-
[28]
Ghosh, C
S. Ghosh, C. K. R. Evuru, S. Kumar, U. Tyagi, O. Nieto, Z. Jin, and D. Manocha. Visual description grounding reduces hallucinations and boosts reasoning in LVLMs.CoRR, abs/2405.15683,
-
[30]
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P . Wang, X. Bi, et al. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning.CoRR, abs/2501.12948,
-
[31]
S. Guo, A. Didolkar, N. R. Ke, A. Goyal, F. Huszár, and B. Schölkopf. Learning beyond pattern matching? assaying mathematical understanding in LLMs.CoRR, abs/2405.15485,
-
[32]
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu. Reasoning with language model is planning with world model. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,
2023
-
[33]
C. He, R. Luo, Y. Bai, S. Hu, Z. L. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, et al. OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems. CoRR, abs/2402.14008, 2024a. M. He, Y. Shen, W. Zhang, Z. Tan, an...
-
[34]
Hendrycks, C
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding.CoRR, abs/2009.03300,
2009 arXiv
-
[35]
G. Hinton. Distilling the knowledge in a neural network.CoRR, abs/1503.02531,
-
[36]
Hosseini, X
A. Hosseini, X. Yuan, N. Malkin, A. Courville, A. Sordoni, and R. Agarwal. V-STaR: Training verifiers for self-taught reasoners.CoRR, abs/2402.06457,
-
[37]
J. Hu. REINFORCE++: A simple and efficient approach for aligning large language models.CoRR, abs/2501.03262,
-
[38]
Huang, S
J. Huang, S. S. Gu, L. Hou, Y. Wu, X. Wang, H. Yu, and J. Han. Large language models can self-improve. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,
2023
-
[39]
Huang, X
Y. Huang, X. Liu, Y. Gong, Z. Gou, Y. Shen, N. Duan, and W. Chen. Key-point-driven data synthesis with its enhancement on mathematical reasoning.CoRR, abs/2403.02333,
-
[40]
C. Jia, P . Wang, Z. Li, Y.-C. Li, Z. Zhang, N. Tang, and Y. Yu. BWArea model: Learning world model, inverse dynamics, and policy for controllable language generation.CoRR, abs/2405.17039,
-
[41]
C. Jia, Z. Li, P . Wang, Y.-C. Li, Z. Hou, Y. Dong, and Y. Yu. Controlling large language model with latent actions.CoRR, abs/2503.21383,
-
[42]
Jie and W
Z. Jie and W. Lu. Leveraging training data in few-shot prompting for numerical reasoning.CoRR, abs/2305.18170,
-
[43]
22 A Survey on Large Language Models for Mathematical Reasoning B. Jin, H. Zeng, Z. Yue, D. Wang, H. Zamani, and J. Han. Search-R1: Training LLMs to reason and leverage search engines with reinforcement learning.CoRR, abs/2503.09516,
-
[44]
A. Joulin. Fasttext.zip: Compressing text classification models.CoRR, abs/1612.03651,
-
[46]
Koncel-Kedziorski, S
R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, and H. Hajishirzi. MAWPS: A math word prob- lem repository. InProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics,
2016
-
[47]
Kwiatkowski, E
T. Kwiatkowski, E. Choi, Y. Artzi, and L. Zettlemoyer. Scaling semantic parsers with on-the-fly ontology matching. InProceedings of the 2013 Conference on Empirical Methods in Natural Language Processing,
2013
-
[49]
Lehnert, S
L. Lehnert, S. Sukhbaatar, D. Su, Q. Zheng, P . Mcvay, M. Rabbat, and Y. Tian. Beyond A*: Better planning with Transformers via search dynamics bootstrapping.CoRR, abs/2402.14083,
-
[50]
Levonian, C
Z. Levonian, C. Li, W. Zhu, A. Gade, O. Henkel, M.-E. Postle, and W. Xing. Retrieval-augmented generation to improve math question-answering: Trade-offs between groundedness and human preference.CoRR, abs/2310.03184,
-
[51]
Lewkowycz, A
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V . Misra. Solving quantitative reasoning problems with language models.CoRR, abs/2206.14858,
-
[52]
C. Li, W. Wang, J. Hu, Y. Wei, N. Zheng, H. Hu, Z. Zhang, and H. Peng. Common 7b language models already possess strong math capabilities.CoRR, abs/2403.04706, 2024a. C. Li, Z. Yuan, H. Yuan, G. Dong, K. Lu, J. Wu, C. Tan, X. Wang, and C. Zhou. MuggleMath: Assessing the impact...
-
[53]
X. Li, H. Zou, and P . Liu. ToRL: Scaling tool-integrated RL.CoRR, abs/2503.23383, 2025a. Y. Li, Z. Lin, S. Zhang, Q. Fu, B. Chen, J.-G. Lou, and W. Chen. Making large language models better reasoners with step-aware verifier.CoRR, abs/2206.02336,
-
[54]
Z. Li, T. Xu, and Y. Yu. When is RL better than DPO in RLHF? a representation and optimization perspective. InThe Second Tiny Papers Track at International Conference on Learning Representations, 2024e. Z. Li, T. Xu, Y. Zhang, Z. Lin, Y. Yu, R. Sun, and Z.-Q. Luo. ReMax: A sim...
-
[55]
Liang, W
Z. Liang, W. Yu, T. Rajpurohit, P . Clark, X. Zhang, and A. Kalyan. Let GPT be a math tutor: Teaching math word problem solvers with customized exercise generation. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,
2023
-
[56]
Lightman, V
H. Lightman, V . Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe. Let’s verify step by step.CoRR, abs/2305.20050,
-
[57]
Y. Lin, S. Seto, M. ter Hoeve, K. Metcalf, B.-J. Theobald, X. Wang, Y. Zhang, C. Huang, and T. Zhang. On the limited generalization capability of the implicit reward model induced by direct preference optimization. CoRR, abs/2409.03650, 2024a. Y.-T. Lin, D. Jin, T. Xu, T. Wu, ...
-
[58]
Z. Lin, Z. Gou, Y. Gong, X. Liu, Y. Shen, R. Xu, C. Lin, Y. Yang, J. Jiao, N. Duan, et al. Rho-1: Not all tokens are what you need.CoRR, abs/2404.07965, 2024b. J. Lindsey, W. Gurnee, E. Ameisen, B. Chen, A. Pearce, N. L. Turner, C. Citro, D. Abrahams, S. Carter, B. Hosmer, J. ...
-
[59]
B. Liu, S. Bubeck, R. Eldan, J. Kulkarni, Y. Li, A. Nguyen, R. Ward, and Y. Zhang. TinyGSM: achieving> 80% on GSM8k with small language models.CoRR, abs/2312.09241,
-
[60]
H. Liu, Y. Zhang, Y. Luo, and A. C. Yao. Augmenting math word problems via iterative question composing. CoRR, abs/2401.09003, 2024a. Z. Liu, H. Liu, D. Zhou, and T. Ma. Chain of thought empowers transformers to solve inherently serial problems. InProceedings of the 12th Inter...
-
[61]
L. Luo, Y. Liu, R. Liu, S. Phatale, H. Lara, Y. Li, L. Shu, Y. Zhu, L. Meng, J. Sun, et al. Improve mathematical reasoning in language models by automated process supervision.CoRR, abs/2406.06592,
-
[62]
Meadows and A
J. Meadows and A. Freitas. A survey in mathematical language processing.CoRR, abs/2205.15231,
-
[63]
Y. Meng, M. Xia, and D. Chen. SimPO: Simple preference optimization with a reference-free reward.CoRR, abs/2405.14734,
-
[64]
Menick, M
J. Menick, M. Trebacz, M. Mikulak-Krupi ´ nski, E. Elsen, A. Valbusa, A. Raichuk, R. Kadlec, G. Irving, J. Kopeck`y, M. Ring, A. Glaese, N. Elhage, T. Hennigan, S. Saci, T. Cai, M. Fritz, A. Jones, D. Pfau, T. Pohlen, and O. Vinyals. Teaching language models to support answers...
-
[65]
Miao, C.-C
24 A Survey on Large Language Models for Mathematical Reasoning S.-Y. Miao, C.-C. Liang, and K.-Y. Su. A diverse corpus for evaluating and developing english math word problem solvers.CoRR, abs/2106.15772,
-
[66]
Mirzadeh, K
S. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M. Farajtabar. GSM-Symbolic: Under- standing the limitations of mathematical reasoning in large language models.CoRR, abs/2410.05229,
-
[67]
Mishra, M
S. Mishra, M. Finlayson, P . Lu, L. Tang, S. Welleck, C. Baral, T. Rajpurohit, O. Tafjord, A. Sabharwal, P . Clark, and A. Kalyan. LILA: A unified benchmark for mathematical reasoning. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,
2022
-
[68]
Muennighoff, Z
N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P . Liang, E. Candès, and T. Hashimoto. s1: Simple test-time scaling.CoRR, abs/2501.19393,
-
[69]
Patel, S
A. Patel, S. Bhattamishra, and N. Goyal. Are NLP models really able to solve simple math word problems? CoRR, abs/2103.07191,
-
[70]
X. Peng, C. Xia, X. Yang, C. Xiong, C.-S. Wu, and C. Xing. ReGenesis: LLMs can grow into reasoning generalists via self-improvement.CoRR, abs/2410.02108,
-
[71]
X. Qu, Y. Li, Z. Su, W. Sun, J. Yan, D. Liu, G. Cui, D. Liu, S. Liang, J. He, P . Li, W. Wei, J. Shao, C. Lu, Y. Zhang, X. Hua, B. Zhou, and Y. Cheng. A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.CoRR, abs/2503.21614,
-
[72]
Schulman, F
J. Schulman, F. Wolski, P . Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. CoRR, abs/1707.06347,
-
[73]
Setlur, C
A. Setlur, C. Nagpal, A. Fisch, X. Geng, J. Eisenstein, R. Agarwal, A. Agarwal, J. Berant, and A. Kumar. Rewarding progress: Scaling automated process verifiers for LLM reasoning.CoRR, abs/2410.08146,
-
[74]
Z. Shao, P . Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models.CoRR, abs/2402.03300,
-
[75]
M. Shen, G. Zeng, Z. Qi, Z. Hong, Z. Chen, W. Lu, G. W. Wornell, S. Das, D. Cox, and C. Gan. Satori: Reinforcement learning with chain-of-action-thought enhances LLM reasoning via autoregressive search. CoRR, abs/2502.02508,
-
[76]
Z. Shen. LLM with tools: A survey.CoRR, abs/2409.18807,
-
[77]
25 A Survey on Large Language Models for Mathematical Reasoning F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, D. Das, and J. Wei. Language models are multilingual chain-of-thought reasoners.CoRR, abs/2210.03057,
-
[79]
Sprague, F
Z. Sprague, F. Yin, J. D. Rodriguez, D. Jiang, M. Wadhwa, P . Singhal, X. Zhao, X. Ye, K. Mahowald, and G. Durrett. To CoT or not to CoT? chain-of-thought helps mainly on math and symbolic reasoning.CoRR, abs/2409.12183, 2024a. Z. Sprague, F. Yin, J. D. Rodriguez, D. Jiang, M....
-
[80]
Stiennon, L
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P . F. Christiano. Learning to summarize from human feedback.CoRR, abs/2009.01325,
2009 arXiv
-
[82]
Q. Tang, Z. Deng, H. Lin, X. Han, Q. Liang, B. Cao, and L. Sun. ToolAlpaca: Generalized tool learning for language models with 3000 simulated cases.CoRR, abs/2306.05301,
-
[83]
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al. Kimi k1.5: Scaling reinforcement learning with LLMs.CoRR, abs/2501.12599,
-
[84]
Toshniwal, W
S. Toshniwal, W. Du, I. Moshkov, B. Kisacanin, A. Ayrapetyan, and I. Gitman. OpenMathInstruct-2: Acceler- ating AI for math with massive open-source instruction data.CoRR, abs/2410.01560, 2024a. S. Toshniwal, I. Moshkov, S. Narenthiran, D. Gitman, F. Jia, and I. Gitman. OpenMa...
-
[86]
B. Wang, S. Min, X. Deng, J. Shen, Y. Wu, L. Zettlemoyer, and H. Sun. Towards understanding chain-of- thought prompting: An empirical study of what matters. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 2023a. E. Z. Wang, F. Cassano...
-
[87]
Wang and D
X. Wang and D. Zhou. Chain-of-thought reasoning without prompting.CoRR, abs/2402.10200,
-
[88]
X. Wang, J. Wei, D. Schuurmans, Q. V . Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou. Self-consistency improves chain of thought reasoning in language models. InProceedings of the 11th International Conference on Learning Representations, 2023c. Y. Wang, X. Liu, and S. S...
2017
-
[89]
Y. Wang, Q. Liu, J. Xu, T. Liang, X. Chen, Z. He, L. Song, D. Yu, J. Li, Z. Zhang, et al. Thoughts are all over the place: On the underthinking of o1-like LLMs.CoRR, abs/2501.18585,
-
[91]
C. Wu, Z. Lin, W. Fang, and Y. Huang. A medical diagnostic assistant based on LLM. InHealth Information Processing Evaluation Track Papers, 2024a. S. Wu, E. M. Shen, C. Badrinath, J. Ma, and H. Lakkaraju. Analyzing chain-of-thought prompting in large language models via gradie...
-
[92]
S. Xu, W. Fu, J. Gao, W. Ye, W. Liu, Z. Mei, G. Wang, C. Yu, and Y. Wu. Is DPO superior to PPO for LLM alignment? a comprehensive study.CoRR, abs/2404.10719,
-
[93]
Y. Yan, J. Su, J. He, F. Fu, X. Zheng, Y. Lyu, K. Wang, S. Wang, Q. Wen, and X. Hu. A survey of mathematical reasoning in the era of multimodal large language model: Benchmark, method & challenges.CoRR, abs/2412.11936,
-
[94]
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Yang, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P . Zhang, P ...
-
[95]
L. Yang, K. Lee, R. Nowak, and D. Papailiopoulos. Looped Transformers are better at learning learning algorithms.CoRR, abs/2311.12424, 2024c. S. Yang, D. Schuurmans, P . Abbeel, and O. Nachum. Chain of thought imitation with procedure cloning. In Advances in Neural Information...
-
[96]
T. Yang, M. Yan, H. Zhao, and T. Yang. LemmaHead: RAG assisted proof generation using large language models.CoRR, abs/2501.15797,
-
[97]
T. Ye, Z. Xu, Y. Li, and Z. Allen-Zhu. Physics of language models: Part 2.2, how to learn from mistakes on grade-school math problems.CoRR, abs/2408.16293,
-
[98]
Y. Ye, Z. Huang, Y. Xiao, E. Chern, S. Xia, and P . Liu. LIMO: Less is more for reasoning.CoRR, abs/2502.03387,
-
[99]
H. Ying, Z. Wu, Y. Geng, J. Wang, D. Lin, and K. Chen. Lean workbook: A large-scale Lean problem set formalized from natural language math problems.CoRR, abs/2406.03847, 2024a. H. Ying, S. Zhang, L. Li, Z. Zhou, Y. Shao, Z. Fei, Y. Ma, J. Hong, K. Liu, Z. Wang, et al. InternLM...
-
[100]
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, et al. DAPO: An open-source LLM reinforcement learning system at scale.CoRR, abs/2503.14476,
-
[101]
L. Yuan, W. Li, H. Chen, G. Cui, N. Ding, K. Zhang, B. Zhou, Z. Liu, and H. Peng. Free process rewards without process labels.CoRR, abs/2412.01981,
-
[102]
28 A Survey on Large Language Models for Mathematical Reasoning X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen. MAmmoTH: Building math generalist models through hybrid instruction tuning. InProceedings of the 12th International Conference on Learning Repr...
-
[103]
Zelikman, G
E. Zelikman, G. Harik, Y. Shao, V . Jayasiri, N. Haber, and N. D. Goodman. Quiet-STaR: Language models can teach themselves to think before speaking.CoRR, abs/2403.09629,
-
[104]
W. Zeng, Y. Huang, Q. Liu, W. Liu, K. He, Z. Ma, and J. He. SimpleRL-Zoo: Investigating and taming zero reinforcement learning for open base models in the wild.CoRR, abs/2503.18892,
-
[105]
Zhang, K
B. Zhang, K. Zhou, X. Wei, X. Zhao, J. Sha, S. Wang, and J. Wen. Evaluating and improving tool-augmented computation-intensive math reasoning. InAdvances in Neural Information Processing Systems 37, 2023a. D. Zhang, L. Wang, L. Zhang, B. T. Dai, and H. T. Shen. The gap of sema...
-
[106]
Zhang, J
D. Zhang, J. Wu, J. Lei, T. Che, J. Li, T. Xie, X. Huang, S. Zhang, M. Pavone, Y. Li, et al. Llama-Berry: Pairwise optimization for o1-like olympiad-level mathematical reasoning.CoRR, abs/2410.02884, 2024a. D. Zhang, S. Zhoubian, Z. Hu, Y. Yue, Y. Dong, and J. Tang. REST-MCTS*...
-
[107]
Zhang, A
L. Zhang, A. Hosseini, H. Bansal, M. Kazemi, A. Kumar, and R. Agarwal. Generative verifiers: Reward modeling as next-token prediction.CoRR, abs/2408.15240, 2024c. Y. Zhang, Y. Luo, Y. Yuan, and A. C. Yao. Autonomous data selection with language models for mathematical texts.Co...
-
[108]
R. Zhao, A. Meterez, S. Kakade, C. Pehlevan, S. Jelassi, and E. Malach. Echo chamber: RL post-training amplifies behaviors learned in pretraining.CoRR, abs/2504.07912,
-
[109]
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al. A survey of large language models.CoRR, abs/2303.18223,
-
[110]
Z. Zhao, H. Dong, A. Saha, C. Xiong, and D. Sahoo. Automatic curriculum expert iteration for reliable LLM reasoning.CoRR, abs/2410.07627,
-
[111]
Zheng, Z
C. Zheng, Z. Zhang, B. Zhang, R. Lin, K. Lu, B. Yu, D. Liu, J. Zhou, and J. Lin. ProcessBench: Identifying process errors in mathematical reasoning.CoRR, abs/2412.06559, 2024a. G. Zheng, B. Yang, J. Tang, H. Zhou, and S. Yang. DDCoT: Duty-distinct chain-of-thought prompting fo...
-
[112]
Zheng, J
K. Zheng, J. M. Han, and S. Polu. MiniF2F: a cross-system benchmark for formal olympiad-level mathematics. CoRR, abs/2109.00110,
-
[113]
Zheng, R
L. Zheng, R. Wang, X. Wang, and B. An. Synapse: Trajectory-as-exemplar prompting with memory for computer control. InProceedings of the 12th International Conference on Learning Representations, 2024b. Q. Zhong, K. Wang, Z. Xu, J. Liu, L. Ding, B. Du, and D. Tao. Achieving >97...
-
[114]
G. Zhou, P . Qiu, C. Chen, J. Wang, Z. Yang, J. Xu, and M. Qiu. Reinforced MLLM: A survey on RL-based reasoning in multimodal large language models.CoRR, abs/2504.21277,
-
[115]
29 A Survey on Large Language Models for Mathematical Reasoning H. Zhou, A. Nova, H. Larochelle, A. C. Courville, B. Neyshabur, and H. Sedghi. Teaching algorithmic reasoning via in-context learning.CoRR, abs/2211.09066,
-
[116]
K. Zhou, B. Zhang, J. Wang, Z. Chen, W. X. Zhao, J. Sha, Z. Sheng, S. Wang, and J.-R. Wen. JiuZhang3.0: Effi- ciently improving mathematical reasoning by training small data synthesis models.CoRR, abs/2405.14365,
-
[117]
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P . Christiano, and G. Irving. Fine- tuning language models from human preferences.CoRR, abs/1909.08593,
1909 arXiv
-
[118]
An essential goal in evaluating mathematical reasoning models is to assess whether they exhibit capabilities comparable to, or exceeding, those of humans
30 A Survey on Large Language Models for Mathematical Reasoning A Mathematical Reasoning Evaluation Benchmarks In this section, we introduce benchmarks that can be used to evaluate a model’s capabilities relevant to mathematical reasoning. An essential goal in evaluating mathe...
2016
-
[119]
[aim, 2024] consists of 30 problems from the 2024 American Mathematics Invitational, designed to evaluate models’ performance on complex mathematical problems. OlympiadBench [He et al., 2024a] is a bilingual, multimodal benchmark designed for Olympiad-level science competition...
2024
-
[1950]
Tutunov, A
R. Tutunov, A. Grosnit, J. Ziomek, J. Wang, and H. Bou-Ammar. Why can large language models generate correct chain-of-thoughts?CoRR, abs/2310.13571,
-
[1962]
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, et al. Emergent abilities of large language models.Transactions on Machine Learning Research, 2022a. J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Ch...
-
[1963]
D. Feng, B. Qin, C. Huang, Z. Zhang, and W. Lei. Towards analyzing and understanding the limitations of DPO: A theoretical perspective.CoRR, abs/2404.04626, 2024a. G. Feng, B. Zhang, Y. Gu, H. Ye, D. He, and L. Wang. Towards revealing the mystery behind Chain of Thought: A the...
-
[1964]
Brown, J
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V . Le, C. Ré, and A. Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling.CoRR, abs/2407.21787,
-
[1965]
Snell, J
C. Snell, J. Lee, K. Xu, and A. Kumar. Scaling LLM test-time compute optimally can be more effective than scaling model parameters.CoRR, abs/2408.03314,
-
[2006]
C. Dai, K. Li, W. Zhou, and S. Hu. Beyond imitation: Learning key reasoning steps from dual chain-of- thoughts in reasoning distillation.CoRR, abs/2405.19737,
-
[2008]
G. Chen, M. Liao, C. Li, and K. Fan. AlphaMath almost zero: Process supervision without process. In Advances in Neural Information Processing Systems 38, 2024a. M. Chen, T. Li, H. Sun, Y. Zhou, C. Zhu, F. Yang, Z. Zhou, W. Chen, H. Wang, J. Z. Pan, et al. Learning to reason wi...
2023 arXiv
-
[2009]
Arora, Y
19 A Survey on Large Language Models for Mathematical Reasoning S. Arora, Y. Li, Y. Liang, T. Ma, and A. Risteski. A latent variable model approach to PMI-based word embeddings.CoRR, abs/1502.03520,
-
[2011]
Gandhi, A
K. Gandhi, A. Chakravarthy, A. Singh, N. Lile, and N. D. Goodman. Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective stars.CoRR, abs/2503.01307,
-
[2012]
E. Y. Chang, Y. Tong, M. Niu, G. Neubig, and X. Yue. Demystifying long chain-of-thought reasoning in LLMs. CoRR, abs/2502.03373,
-
[2013]
X. Lai, Z. Tian, Y. Chen, S. Yang, X. Peng, and J. Jia. Step-DPO: Step-wise preference optimization for long-chain reasoning of LLMs.CoRR, abs/2406.18629,
-
[2014]
X. Guan, L. L. Zhang, Y. Liu, N. Shang, Y. Sun, Y. Zhu, F. Yang, and M. Yang. rStar-Math: Small LLMs can master math reasoning with self-evolved deep thinking.CoRR, abs/2501.04519,
-
[2016]
Kazemnejad, M
A. Kazemnejad, M. Aghajohari, E. Portelance, A. Sordoni, S. Reddy, A. Courville, and N. L. Roux. VinePPO: Unlocking RL potential for LLM reasoning through refined credit assignment.CoRR, abs/2410.01679,
-
[2017]
Arcuschin, J
I. Arcuschin, J. Janiak, R. Krzyzanowski, S. Rajamanoharan, N. Nanda, and A. Conmy. Chain-of-Thought reasoning in the wild is not always faithful.CoRR, abs/2503.08679,
-
[2019]
S. An, Z. Ma, Z. Lin, N. Zheng, J.-G. Lou, and W. Chen. Learning from mistakes makes LLM better reasoner. CoRR, abs/2310.20689,
-
[2020]
Sui, Y.-N
Y. Sui, Y.-N. Chuang, G. Wang, J. Zhang, T. Zhang, J. Yuan, H. Liu, A. Wen, H. Chen, X. Hu, et al. Stop overthinking: A survey on efficient reasoning for large language models.CoRR, abs/2503.16419,
-
[2021]
Austin, A
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al. Program synthesis with large language models.CoRR, abs/2108.07732,
-
[2022]
Bansal, A
H. Bansal, A. Hosseini, R. Agarwal, V . Q. Tran, and M. Kazemi. Smaller, weaker, yet better: Training LLM reasoners via compute-optimal sampling.CoRR, abs/2408.16737,
-
[2023]
Ankner, M
Z. Ankner, M. Paul, B. Cui, J. D. Chang, and P . Ammanabrolu. Critique-out-loud reward models.CoRR, abs/2408.11791,
-
[2024]
Mathematical Association of America. S. Adarsh, K. Shridhar, C. Gulcehre, N. Monath, and M. Sachan. SIKeD: Self-guided iterative knowledge distillation for mathematical reasoning.CoRR, abs/2410.18574,
-
[2025]
B. Gao, Z. Cai, R. Xu, P . Wang, C. Zheng, R. Lin, K. Lu, J. Lin, C. Zhou, W. Xiao, J. Hu, T. Liu, and B. Chang. LLM critics help catch bugs in mathematics: Towards a better mathematical verifier with natural language feedback.CoRR, abs/2406.14024,
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.