Pith. sign in

REVIEW 29 cited by

Large Language Models for Mathematical Reasoning: Progresses and Challenges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.00157 v4 pith:PJ7KVVZ2 submitted 2024-01-31 cs.CL

classification cs.CL
keywords mathematicalbeenchallengesllmsdatasetsfieldlandscapelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mathematical reasoning serves as a cornerstone for assessing the fundamental cognitive capabilities of human intelligence. In recent times, there has been a notable surge in the development of Large Language Models (LLMs) geared towards the automated resolution of mathematical problems. However, the landscape of mathematical problem types is vast and varied, with LLM-oriented techniques undergoing evaluation across diverse datasets and settings. This diversity makes it challenging to discern the true advancements and obstacles within this burgeoning field. This survey endeavors to address four pivotal dimensions: i) a comprehensive exploration of the various mathematical problems and their corresponding datasets that have been investigated; ii) an examination of the spectrum of LLM-oriented techniques that have been proposed for mathematical problem-solving; iii) an overview of factors and concerns affecting LLMs in solving math; and iv) an elucidation of the persisting challenges within this domain. To the best of our knowledge, this survey stands as one of the first extensive examinations of the landscape of LLMs in the realm of mathematics, providing a holistic perspective on the current state, accomplishments, and future challenges in this rapidly evolving field.

Discussion (0). Sign in to comment.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Superloop Equations and Minimal Surfaces I: Confining minimal surface in $4D, N=1$ SYM

    hep-th 2026-08 conditional novelty 7.0 of 10

    A geometrically constructed surface-area phase is proven to dress any solution of the finite-N N=1 SYM superloop hierarchy and produces a rectangular Wilson phase exp(-iσLT) with arbitrary positive σ.

  2. Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A one-step semiparametric estimator using pairwise comparison signals as control variates achieves the efficiency bound for estimating LLM accuracy on math benchmarks.

  3. Adaptive Information Control for Search-Augmented LLM Reasoning

    cs.CL 2026-02 conditional novelty 6.0 of 10

    DeepControl uses information-utility signals to control when search-augmented reasoning agents stop retrieving and how much evidence they expand, improving QA accuracy across seven benchmarks and two model sizes.

  4. One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

    cs.SE 2025-12 conditional novelty 6.0 of 10

    Repository-level issue localization can be done by a single jump-to-definition tool trained with reinforcement learning, achieving strong results on SWE-bench despite using only open-weights models.

  5. Arrows of Math Reasoning Data Synthesis for Large Language Models: Diversity, Complexity and Correctness

    cs.CL 2025-08 reject novelty 6.0 of 10

    A program-assisted pipeline generates 12.3 million math problem-solution pairs with execution-based verification, and fine-tuning on a 50k sample improves model scores on GSM8K, MATH, Minerva, and SVAMP.

  6. GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A generative multimodal process reward model that produces step-level critiques and corrections improves average math accuracy for six multimodal LLMs by 2.9 to 5.9 points under a refinement-based Best-of-N strategy.

  7. WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new wargame-based benchmark finds large language models score far below human experts on strategic reasoning across situation awareness, opponent modeling, and policy generation.

  8. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Reasoning fine-tuning makes LLMs more accurate on answerable problems but worse at abstaining on unanswerable ones, across a new 20-dataset benchmark.

  9. Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Safe uses step-level formal verification in Lean 4, aggregated by a small LSTM and combined with process reward scores, to improve best-of-n accuracy for LLM mathematical reasoning.

  10. Structured Pruning for Diverse Best-of-N Reasoning Optimization

    cs.CL 2025-06 reject novelty 6.0 of 10

    SPRINT learns to select which attention heads to prune per question, improving Pass@N over random head selection and multinomial sampling on MATH500 and GSM8K.

  11. More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Comparative words in prompts can shift LLM answers toward the framed direction in simple arithmetic comparisons, with demographic terms amplifying the effect.

  12. Learning to Insert [PAUSE] Tokens for Better Reasoning

    cs.CL 2025-06 reject novelty 6.0 of 10

    A likelihood-based [PAUSE] token insertion method for fine-tuning shows small gains on GSM8K and MBPP, but the AQUA-RAT result is unreliable because the test set contains training samples.

  13. Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLM reasoning can be scored separately for knowledge and step-by-step information gain, and doing so shows SFT and RL affect these two capacities differently across medicine and math.

  14. Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    PCPO selects preference pairs by combining correct-answer status with token-level probability consistency, then trains with a weighted DPO+NLL loss, yielding small and partly inconsistent gains over outcome-only metho...

  15. CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...

  16. Can reasoning models comprehend mathematical problems in Chinese ancient texts? An empirical study based on data from Suanjing Shishu

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Reasoning LLMs partially solve classical Chinese math problems from Suanjing Shishu, reaching up to 70% accuracy with original solution methods provided, but lag behind their modern-math performance.

  17. Assessing the Benefits of Combining Advanced Deep Learning Techniques for Post-Disaster Building Damage Assessment from UAV Imagery

    cs.CV 2026-08 conditional novelty 5.0 of 10

    A hybrid CV+LVLM pipeline improves post-disaster building damage counting over single models in some configurations, but fails in others and shows low absolute accuracy.

  18. STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Compressing multi-hop search trajectories into per-answer evidence cards improves final answer selection over raw-trajectory or string-only comparison on four multi-hop QA benchmarks.

  19. Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving

    cs.AI 2026-07 conditional novelty 5.0 of 10

    LLM math accuracy varies across equivalent problem representations, and executable-reasoning scaffolding redistributes rather than removes the errors.

  20. ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

    cs.AI 2025-12 reject novelty 5.0 of 10

    LLM reasoning benchmark scores vary substantially across repeated runs under the same model, strategy, and task, so single-run evaluation can misrank systems.

  21. Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding

    cs.CL 2025-07 conditional novelty 5.0 of 10

    DeepSeek-R1-distilled models show higher multi-document QA accuracy than their base counterparts and flatter position-bias curves, especially with 50-80 documents.

  22. Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Across ten LLMs, masking the final answer inside a complete reasoning chain causes a 26.9-point accuracy drop, evidence that models anchor to answers, not reasoning templates.

  23. Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A new benchmark (POBs) reveals that LLMs lean progressive-collectivist, that test-time compute offers limited gains in neutrality or consistency, and that newer model versions often become more biased and less consistent.

  24. Towards General Continuous Memory for Vision-Language Models

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A vision-language model can act as its own continuous memory encoder, compressing external multimodal knowledge into eight embeddings that improve reasoning when prepended to the frozen model.

  25. Feature Generation Using LLMs: An Evolutionary Algorithm Approach

    cs.LG 2026-06 conditional novelty 4.0 of 10

    A funsearch-style evolutionary loop using LLaMA-3.1 7B-generated Python expressions creates new table features and improves F1 in 13 of 16 evaluated classification settings.

  26. From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning

    cs.AI 2026-01 conditional novelty 4.0 of 10

    Post-training LLMs first on abstract number-free reasoning plans (CoMT), then with confidence-weighted rewards (CCRL), raises math accuracy by ~2-5 points over standard CoT-SFT+RL across four models.

  27. A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis

    cs.CL 2025-06 conditional novelty 4.0 of 10

    An LLM agent that reframes beam analysis as OpenSeesPy code generation reaches over 99 percent reliability on a small benchmark, but chiefly because the prompt contains a near-identical solved example.

  28. WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis

    cs.CY 2025-06 conditional novelty 4.0 of 10

    A GPT-4o-based smart tutor with problem-specific documents provided homework feedback in a circuit analysis course, and 90.9% of 66 student feedback responses were positive.

  29. Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration

    eess.SP 2025-06 conditional novelty 4.0 of 10

    The paper proposes a systematic classification and two roadmaps for using foundation models (LLMs and wireless foundation models) to design Synesthesia of Machines systems for 6G, with preliminary case-study evidence ...

Pith tools