REVIEW 8 cited by
GPT Can Solve Mathematical Problems Without a Calculator
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Previous studies have typically assumed that large language models are unable to accurately perform arithmetic operations, particularly multiplication of >8 digits, and operations involving decimals and fractions, without the use of calculator tools. This paper aims to challenge this misconception. With sufficient training data, a 2 billion-parameter language model can accurately perform multi-digit arithmetic operations with almost 100% accuracy without data leakage, significantly surpassing GPT-4 (whose multi-digit multiplication accuracy is only 4.3%). We also demonstrate that our MathGLM, fine-tuned from GLM-10B on a dataset with additional multi-step arithmetic operations and math problems described in text, achieves similar performance to GPT-4 on a 5,000-samples Chinese math problem test set. Our code and data are public at https://github.com/THUDM/MathGLM.
Forward citations
Cited by 8 Pith papers
-
Self-guided Knowledgeable Network of Thoughts: Amplifying Reasoning with Large Language Models
A large language model can plan and write an executable network of elementary reasoning steps for itself, and this self-guided workflow beats prior prompting schemes on sorting, arithmetic, and language counting tasks.
-
MultiLingPoT: Enhancing Mathematical Reasoning with Multilingual Program Fine-tuning
Training one model on Program-of-Thought solutions in four programming languages raises math accuracy for each language, and answer mixing outperforms single-language augmented training by up to about 6 percentage points.
-
Mars-PO: Multi-Agent Reasoning System Preference Optimization
A multi-agent preference optimization method that uses pooled correct answers from several LLMs as shared positives and each model's own errors as negatives improves math reasoning accuracy on GSM8K and MATH.
-
Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies
A small transformer learns addition, multiplication, and division by mastering simple digit subtasks first, and human teaching strategies lift its arithmetic accuracy to ~100%.
-
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
Tool-augmented LLM annotators improve agreement with ground-truth preferences on long-form factual and coding tasks, with mixed results on math, compared to standard LLM-as-a-judge baselines.
-
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
PerPO trains multimodal LLMs by ranking their candidate answers with deterministic visual rewards (IoU, edit distance) and using the reward differences as margins in listwise preference optimization.
-
Perspective Transition of Large Language Models for Solving Subjective Tasks
Reasoning through Perspective Transition (RPT) improves LLM performance on subjective NLP tasks by ranking direct, role, and third-person perspectives by self-reported confidence and answering from the top-ranked perspective.
-
Scaling Particle Collision Data Analysis
A byte-level transformer trained from scratch on simulated collider data matches specialized jet-tagging models when given enough training examples.
Discussion (0). Continue with ORCID to comment.