Benchmarks in Leipzig

Alejandro Morales; Alexander Ivanov; Alexander Taveira Blomenhofer; Andrea Rosana; Andrei Balakin; Annette Werner; Aryaman Jal; Baran Hashemi; Bernd Sturmfels; Carl Felix Waller

arxiv: 2606.05818 · v1 · pith:7UY2KIUMnew · submitted 2026-06-04 · 🧮 math.HO · cs.AI· math.AG· math.CO· math.RT

Benchmarks in Leipzig

Andrei Balakin , Mikl\'os B\'ona , Marie-Charlotte Brandenburg , Clara Briand , Veronica Calvo Cortes , Shelby Cox , Jesus A. De Loera , Danai Deligeorgaki

show 40 more authors

Hannah Friedman Tim Gehrunger Chiara Giardino Stephen Griffeth Baran Hashemi Elena Hoster Alexander Ivanov Nupur Jain Aryaman Jal Leonie Kayser Joris Koefler Kevin K\"uhn Mario Kummer Felix Lotter Ren\'e Marczinzik Victor S. Miller Alejandro Morales Greta Panova Gianni Petrella Nathan Pflueger Lakshmi Ramesh Nikolas Rieke Carlos Rodriguez Andrea Rosana Flavio Salizzoni Otto T.P. Schmidt Sven Ulf Schmitz Lina Maria Simbaqueba Marin Luca Sodomaco Christian Stump Bernd Sturmfels Alexander Taveira Blomenhofer Simon Telen Philipp Tuchel Emil Verkama Carl Felix Waller Julian Weigert Annette Werner Nathan Williams Claudius Zibrowius

This is my paper

classification 🧮 math.HO cs.AImath.AGmath.COmath.RT

keywords questionsleipzigstageattemptbenchmarksllmsmathematicsmodels

0 comments

read the original abstract

Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3-day workshop *Benchmarks in Leipzig* with 35 participants at the Max Planck Institute for Mathematics in the Sciences in Leipzig, Germany. We present the resulting collection of 100 questions. We evaluated these questions in three stages: a single attempt by five state-of-the-art LLMs, followed by a 20-runs-per-model evaluation with three of these models, and finally a 3-run attempt with two heavy-thinking models. After Stage 1, 41 questions remained completely unsolved; after Stage 2, this count dropped to 16; and we concluded Stage 3 with only 2 unsolved questions. This demonstrates that the mathematical reasoning capabilities of LLMs are becoming impressive.

This paper has not been read by Pith yet.

Benchmarks in Leipzig

discussion (0)