Pith. sign in

REVIEW 15 cited by

Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.21934 v5 pith:WK7MJNIP submitted 2025-03-27 cs.CL

classification cs.CL
keywords reasoningmodelsmathematicalllmsproofachievebenchmarksgemini-2
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent math benchmarks for large language models (LLMs) such as MathArena indicate that state-of-the-art reasoning models achieve impressive performance on mathematical competitions like AIME, with the leading model, Gemini-2.5-Pro, achieving scores comparable to top human competitors. However, these benchmarks evaluate models solely based on final numerical answers, neglecting rigorous reasoning and proof generation which are essential for real-world mathematical tasks. To address this, we introduce a comprehensive evaluation of full-solution reasoning for challenging mathematical problems. Using expert human annotators, we evaluated several state-of-the-art reasoning models on the six problems from the 2025 USAMO within hours of their release. Our results reveal that all tested models struggled significantly: only Gemini-2.5-Pro achieves a non-trivial score of 25%, while all other models achieve less than 5%. Through detailed analysis of reasoning traces, we identify the most common failure modes and find several unwanted artifacts arising from the optimization strategies employed during model training. Overall, our results suggest that current LLMs are inadequate for rigorous mathematical reasoning tasks, highlighting the need for substantial improvements in reasoning and proof generation capabilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cost-Effective Automated Judging of Natural-Language Mathematical Proofs

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Cheap open-weight judges match frontier models on pass/fail grading of IMO proofs at up to 100x lower cost, with the best voting rule still pending replication.

  2. QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

    cs.LG 2026-02 conditional novelty 6.0 of 10

    On QEDBench, frontier LLM judges over-score university math proofs by up to +0.36 on average relative to human experts, while some solver models fail badly on discrete-combinatorial problems.

  3. Large Lemma Miners: Can LLMs do Induction Proofs for Hardware?

    cs.LO 2025-11 conditional novelty 6.0 of 10

    LLMs, verified by a symbolic model checker, produced correct inductive strengthenings for 82 of 94 curated RTL safety properties.

  4. Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

    cs.LG 2025-10 conditional novelty 6.0 of 10

    A dynamic replay and reweighting scheduler (RECAP) preserves general capabilities during RLVR while keeping reasoning performance at least as good as reasoning-only finetuning.

  5. Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling

    cs.LG 2025-10 conditional novelty 6.0 of 10

    When a cosine schedule would halve the learning rate, Seesaw cuts it by √2 and doubles the batch, matching loss curves with ~36% fewer serial steps.

  6. SSRL: Self-Search Reinforcement Learning

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    SSRL, a training pipeline that uses an LLM's own repeated sampling as a search environment for RL, improves question answering without external tools and transfers to real search engines.

  7. Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Seed-Prover and Seed-Geometry prove 121 of 155 formalized past IMO problems, reach 99.6% on MiniF2F-test, and solve 5 of 6 IMO 2025 problems after the competition deadline.

  8. Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    An inference-only neuro-symbolic pipeline, DSP+, solves 80.7% of miniF2F and the previously unsolved imo_2019_p1, matching heavily RL-trained theorem provers without fine-tuning.

  9. Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

    cs.LG 2026-08 conditional novelty 5.0 of 10

    A self-evolving curriculum that retrains a language model on variants of problems it can mostly get right lifts AIME pass@1 from 5.6% to 16.5%, beating static augmentation under the same data budget.

  10. Too long; didn't solve

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    Prompt length and solution length both rise with LLM failure on expert-authored adversarial math problems, linking structural length to empirical difficulty.

  11. No Free Lunch: Rethinking Internal Feedback for LLM Reasoning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Internal feedback rewards (entropy and self-certainty) improve base LLM math reasoning only in early training and degrade later, with little benefit for instruct models.

  12. Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning

    cs.CL 2025-06 conditional novelty 5.0 of 10

    On a new Olympiad-style benchmark, outcome-rewarded LLMs reach 80.1% answer accuracy but only 39.7% reasoning correctness, and a step-by-step LLM verifier outperforms holistic judges at spotting flawed reasoning.

  13. Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems

    cs.AI 2025-05 conditional novelty 5.0 of 10

    LLM-generated search heuristics run through the CPro1 protocol with the reasoning model o3-mini-high produced verified constructions resolving open instances in 7 Handbook design families and newer problems.

  14. The Mathematician's Assistant: Integrating AI into Research Practice

    math.HO 2025-08 conditional novelty 3.0 of 10

    AI tools currently augment rather than replace mathematicians, and a five-principle, seven-area framework can help researchers use them responsibly.

  15. On the Surprising Efficacy of LLMs for Penetration-Testing

    cs.CR 2025-07 conditional novelty 3.0 of 10

    A critical review arguing that LLMs are surprisingly effective for penetration testing because the task is largely pattern-matching, while noting serious reliability, safety, and cost barriers to autonomous use.

Pith tools