Pith. sign in

REVIEW 8 cited by

GPT Can Solve Mathematical Problems Without a Calculator

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.03241 v2 pith:NNEZMSEV submitted 2023-09-06 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords operationsarithmeticdatawithoutaccuracyaccuratelycalculatorgpt-4
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Previous studies have typically assumed that large language models are unable to accurately perform arithmetic operations, particularly multiplication of >8 digits, and operations involving decimals and fractions, without the use of calculator tools. This paper aims to challenge this misconception. With sufficient training data, a 2 billion-parameter language model can accurately perform multi-digit arithmetic operations with almost 100% accuracy without data leakage, significantly surpassing GPT-4 (whose multi-digit multiplication accuracy is only 4.3%). We also demonstrate that our MathGLM, fine-tuned from GLM-10B on a dataset with additional multi-step arithmetic operations and math problems described in text, achieves similar performance to GPT-4 on a 5,000-samples Chinese math problem test set. Our code and data are public at https://github.com/THUDM/MathGLM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 10 citations worldwide. Full citation record

  1. Self-guided Knowledgeable Network of Thoughts: Amplifying Reasoning with Large Language Models

    cs.MA 2024-12 conditional novelty 6.0 of 10

    A large language model can plan and write an executable network of elementary reasoning steps for itself, and this self-guided workflow beats prior prompting schemes on sorting, arithmetic, and language counting tasks.

  2. MultiLingPoT: Enhancing Mathematical Reasoning with Multilingual Program Fine-tuning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Training one model on Program-of-Thought solutions in four programming languages raises math accuracy for each language, and answer mixing outperforms single-language augmented training by up to about 6 percentage points.

  3. Mars-PO: Multi-Agent Reasoning System Preference Optimization

    cs.AI 2024-11 conditional novelty 6.0 of 10

    A multi-agent preference optimization method that uses pooled correct answers from several LLMs as shared positives and each model's own errors as negatives improves math reasoning accuracy on GSM8K and MATH.

  4. Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A small transformer learns addition, multiplication, and division by mastering simple digit subtasks first, and human teaching strategies lift its arithmetic accuracy to ~100%.

  5. Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Tool-augmented LLM annotators improve agreement with ground-truth preferences on long-form factual and coding tasks, with mixed results on math, compared to standard LLM-as-a-judge baselines.

  6. PerPO: Perceptual Preference Optimization via Discriminative Rewarding

    cs.AI 2025-02 conditional novelty 5.0 of 10

    PerPO trains multimodal LLMs by ranking their candidate answers with deterministic visual rewards (IoU, edit distance) and using the reward differences as margins in listwise preference optimization.

  7. Perspective Transition of Large Language Models for Solving Subjective Tasks

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Reasoning through Perspective Transition (RPT) improves LLM performance on subjective NLP tasks by ranking direct, role, and third-person perspectives by self-reported confidence and answering from the top-ranked perspective.

  8. Scaling Particle Collision Data Analysis

    cs.LG 2024-11 conditional novelty 4.0 of 10

    A byte-level transformer trained from scratch on simulated collider data matches specialized jet-tagging models when given enough training examples.

Pith tools