Pith. sign in

REVIEW 3 cited by

CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06961 v2 pith:6K5QI4RB submitted 2024-01-13 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords modelshintsproblemsreasoningconceptsinformationllmsmath
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent large language models (LLMs) have shown indications of mathematical reasoning ability on challenging competition-level problems, especially with self-generated verbalizations of intermediate reasoning steps (i.e., chain-of-thought prompting). However, current evaluations mainly focus on the end-to-end final answer correctness, and it is unclear whether LLMs can make use of helpful side information such as problem-specific hints. In this paper, we propose a challenging benchmark dataset for enabling such analyses. The Concept and Hint-Annotated Math Problems (CHAMP) consists of high school math competition problems, annotated with concepts, or general math facts, and hints, or problem-specific tricks. These annotations allow us to explore the effects of additional information, such as relevant hints, misleading concepts, or related problems. This benchmark is difficult, with the best model only scoring 58.1% in standard settings. With concepts and hints, performance sometimes improves, indicating that some models can make use of such side information. Furthermore, we annotate model-generated solutions for their correctness. Using this corpus, we find that models often arrive at the correct final answer through wrong reasoning steps. In addition, we test whether models are able to verify these solutions, and find that most models struggle.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A two-track Olympiad benchmark with deterministic integer answers and step-by-step proof grading shows frontier LLMs drop sharply versus older math benchmarks.

  2. ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark of perturbed symbolic math problems shows that large language models' performance drops sharply under minor numeric, symbolic, and equivalence transformations.

  3. Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device Agents

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Division-of-Thoughts reduces LLM API cost and latency by about 84% and 66% on seven reasoning benchmarks via subtask decomposition and small/large model routing, with accuracy near cloud-only baselines.

Pith tools