REVIEW 3 major objections 4 minor 44 references
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Even the strongest multimodal model fails more than half of image-based Kangaroo math questions, while doing much better on text-only ones, evidence that diagrams are underused.
desk verdict Useful new multilingual visual math benchmark and model snapshot, but the underutilization-of-diagrams conclusion is confounded by uncontrolled question sets and should be reframed as a hypothesis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the multilingual Kangaroo dataset: contest questions from 2014 to 2024 in four languages, each annotated with whether it contains an image, categorized by mathematical topic, and presented to models in a structured prompt that asks for reasoning before the final answer. The key comparison is accuracy on image-based versus non-image questions, which the authors treat as a probe for whether models use diagrammatic information. A co-occurrence heatmap analysis and an analysis of answers on fresh contest questions from the Comunitat Valenciana are used to test whether models reason rather than recite.
What would settle it
Present the same Kangaroo problem in two forms, one with the original diagram and one with the diagram's essential information restated in text, and check whether models' accuracy changes; if models do not lose accuracy when the diagram is removed or made decorative, the claim that they underutilize diagrams is falsified, and if the text-only versions of the same problems are easier, then the benchmark's image/text gap reflects question difficulty rather than diagram use.
Extended reading notes
Core claim
The core discovery is a capability ranking and a behavioral pattern: on image-based Kangaroo questions, Gemini 2.0 Flash achieves the highest precision (45.4%), followed by Qwen-VL 2.5 72B (43.5%) and GPT-4o (40.2%), while most models shift to higher accuracy on text-only questions (Gemini 2.0 Flash 75.9%, Qwen-VL 2.5 72B 70.6%, GPT-4o 65.3%). The authors interpret the accuracy gap as indicating that models underutilize diagrammatic information rather than being unable to reason about the underlying mathematics. No model excels across all mathematical topics (geometry, visual algebra, logic, patterns, combinatorics), and precision varies by language and difficulty, with visual logic and reasoning being the hardest category for all models.
Load-bearing premise
The central comparison assumes that questions with images and questions without images are equally hard, so that any accuracy gap must come from how models handle diagrams rather than from differences in the questions themselves.
Editorial extensions
If this is right
- If the accuracy gap between image and text questions does reflect underutilization of diagrams, then current MLLMs are leaving a large fraction of visual math performance on the table, and training that teaches models to use diagrams should be prioritized over pure parameter scaling.
- The benchmark gives future models a concrete target: reaching human-level performance on Kangaroo image questions requires improving on the current best image-based precision of 45.4% by a wide margin, with the hardest categories being visual logic and reasoning.
- Because no model leads on every mathematical topic, in-domain specialization is not yet possible; the topic-level heatmap suggests each family has specific strengths, but the co-occurrence analysis shows that correlated failures will limit simple ensemble strategies.
- The strong performance of Gemini 2.0 Flash and Qwen-VL 2.5 72B on text-only questions, combined with their bigger drops on image questions, indicates that multilingual textual math reasoning is already strong while visual grounding is the bottleneck.
Reading between the lines
- Editorial inference: The clean way to test the underutilization claim is to compare matched versions of the same problem with and without the diagram; the paper's current design compares different questions, so the language and difficulty mix can explain part of the gap.
- Editorial inference: The benchmark could be extended to measure diagram comprehension directly, for instance by asking models to read off marked angles, lengths, or counts from figures before solving, separating perception errors from reasoning errors.
- Editorial inference: The near-human performance gap and the topic-specific leaders suggest that progress may come less from scaling and more from training that forces models to ground each reasoning step in the image, possibly using intermediate structured representations of the diagram.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a multilingual visual-mathematics benchmark derived from Kangaroo Mathematics Competition (KMC) tests in English, French, Spanish, and Catalan, and evaluates nine multimodal LLMs on accuracy for image-based and text-only questions. The authors report four findings: (1) no model excels across all mathematical topics; (2) most models perform better on questions without images, which is interpreted as underutilization of diagrammatic information; (3) substantial variation across languages and difficulty levels, with Gemini 2.0 Flash achieving the highest precision on image-based tasks; and (4) a qualitative analysis suggesting Gemini and GPT-4o engage in structured reasoning while Pixtral and Llama often fall back on heuristics or random choice. The dataset and code are released on GitHub.
Significance. If the comparisons were properly controlled, the paper would provide a useful multilingual, diagram-focused benchmark for MLLM mathematical reasoning, with a current capability ranking across nine models and a released dataset that can support further research. The authors' choice of KMC tests, the four-language setup, and the topic categorization for image-based questions are strengths, as is the availability of the data and code. However, the headline inference about underutilization of diagrams rests on an uncontrolled comparison between different question sets, so the significance of the main interpretative claim is currently limited.
major comments (3)
- [Section 4, Table 4, Figure 1] The comparison between image-based and non-image accuracy is computed on different KMC question sets, not on matched items or on the same questions rendered with and without images. The abstract and Section 5 interpret the resulting accuracy gap as evidence that models 'underutilize diagrammatic and visual information,' but this conclusion does not follow from the reported design. The authors themselves note in Section 4 that the upward accuracy trend with difficulty level 'may be linked to the reduced presence of images in higher-level questions,' which means image presence is correlated with difficulty. Without a matched control for difficulty and topic mix, the observed text-over-image gain could simply reflect that the text-only questions are easier or have a different topic composition. I recommend either restricting the claim to a descriptive observation about accuracy on two different subsets, or adding a controlled comparison (for example, text-only renderings of the same image-based questions, or per-difficulty/per-topic stratification).
- [Figure 2 and language-level comparisons] The Figure 2 caption states that models are 'not evaluated on the same questions in each language.' Consequently, cross-language accuracy differences and difficulty-level comparisons are confounded by the fact that each language version of the KMC contains different items. The third finding in the abstract, 'substantial variation exists across languages and difficulty levels,' is therefore only a descriptive statement about these particular test forms; it does not establish that the models' multilingual ability differs. To support the language-variation claim, the authors would need per-item translated equivalents or an explicit analysis of item difficulty across languages.
- [Section 5, recitation analysis] The analysis aimed at distinguishing reasoning from recitation is reported qualitatively and without quantitative coding criteria or inter-rater reliability. The text states that Pixtral and Llama 'frequently returned No answer responses' and that Gemini and GPT-4o demonstrated 'more coherent and structured reasoning,' but no counts, examples, or error taxonomy are provided. As written, the fourth abstract finding is not supported by the reported evidence. I suggest adding a small quantitative coding scheme for reasoning quality, or explicitly labeling this part as anecdotal and removing it from the abstract's list of findings.
minor comments (4)
- [Section 4, first paragraph] The word 'Kangarooo' appears to be a typo and should read 'Kangaroo.'
- [Introduction, last paragraph before Section 2] The phrase 'an escellent text to check for recitation' should be corrected to 'an excellent test to check for recitation.'
- [Table 1, caption] The 'Results (%)' column is described as showing image-based and non-image results on the left and right sides, but the caption could state this explicitly in the table header rather than only in the caption text.
- [Figure 3, caption] The sentence 'The figure should be interpreted separately with non-normalized values' is unclear; please specify which values are non-normalized and why the heatmaps must be read separately.
Circularity Check
No significant circularity: the reported rankings and findings are independent measurements on external KMC contest data.
full rationale
This is an empirical benchmark evaluation, not a derivation. The dataset is compiled from public Kangaroo Mathematics Competition tests (2014–2024) in four languages, and the headline accuracies are obtained by prompting nine MLLMs against official ground truth. No parameter is fitted to the reported results, no quantity is defined in terms of the quantity it is said to predict, and no uniqueness theorem or ansatz is imported from the authors' prior work. The only self-citation ([14], a prior text-only Kangaroo study by the same group) appears in the introduction as motivation ('We have recently investigated the ability of LLMs augmented with dedicated reasoners... [14]') and is not load-bearing for any of the four key findings. The main methodological caveat — the abstract's inference that text-over-image gains indicate 'underutilization of diagrammatic information' — rests on comparing different question sets, and the Figure 2 caption itself notes that models 'are not evaluated on the same questions in each language.' That is a validity/confound concern about the strength of the inference, not a self-referential or definitional reduction. Similarly, the reasoning-vs-recitation analysis is a qualitative design choice and not a circular step. Because the central claims are grounded in external, public contest data and independent model evaluations, the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption KMC questions with attached figures require diagram understanding to solve correctly.
- domain assumption The 2014-2024 KMC tests compiled from different countries and languages are treated as comparable difficulty levels despite different question sets.
- domain assumption The Comunitat Valenciana test administered in early April is unseen by the models, so performance on it isolates reasoning from recitation.
Cite this review
Pith. "Pith review of Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests." pith.science (2026). https://pith.science/paper/LTQFTQ46
@misc{pith2026250607418,
author = {Pith},
title = {Pith review of: Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests},
year = {2026},
howpublished = {\url{https://pith.science/paper/LTQFTQ46}},
note = {Machine review of arXiv:2506.07418}
}
read the original abstract
Multimodal Large Language Models (MLLMs) promise advanced vision language capabilities, yet their effectiveness in visually presented mathematics remains underexplored. This paper analyzes the development and evaluation of MLLMs for mathematical problem solving, focusing on diagrams, multilingual text, and symbolic notation. We then assess several models, including GPT 4o, Pixtral, Qwen VL, Llama 3.2 Vision variants, and Gemini 2.0 Flash in a multilingual Kangaroo style benchmark spanning English, French, Spanish, and Catalan. Our experiments reveal four key findings. First, overall precision remains moderate across geometry, visual algebra, logic, patterns, and combinatorics: no single model excels in every topic. Second, while most models see improved accuracy with questions that do not have images, the gain is often limited; performance for some remains nearly unchanged without visual input, indicating underutilization of diagrammatic information. Third, substantial variation exists across languages and difficulty levels: models frequently handle easier items but struggle with advanced geometry and combinatorial reasoning. Notably, Gemini 2.0 Flash achieves the highest precision on image based tasks, followed by Qwen VL 2.5 72B and GPT 4o, though none approach human level performance. Fourth, a complementary analysis aimed at distinguishing whether models reason or simply recite reveals that Gemini and GPT 4o stand out for their structured reasoning and consistent accuracy. In contrast, Pixtral and Llama exhibit less consistent reasoning, often defaulting to heuristics or randomness when unable to align their outputs with the given answer options.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
- [1]
- [2]
- [3]
-
[4]
H. Liu, C. Li, Q. Wu, and Y. J. Lee. LLaV A: Large language and vision assistant.arXiv preprint, arXiv:2304.08485, 2023
arXiv 2023
-
[5]
T. Andreescu and R. Gelca.Mathematical Olympiad Challenges. Springer Science & Business Media, 2008
work page 2008
- [6]
-
[7]
S. Peng, D. Fu, L. Gao, X. Zhong, H. Fu, and Z. Tang. MultiMath: Bridging visual and mathematical reasoning for Large Language Models.arXiv preprint, arXiv:2409.00147, 2024
arXiv 2024
-
[8]
W. Shi, Z. Hu, Y. Bin, J. Liu, Y. Yang, S.-K. Ng, L. Bing, and R. K.-W. Lee. Math-llava: Bootstrapping mathematical reasoning for Multimodal Large Language Models.arXiv preprint arXiv:2406.17294, 2024
arXiv 2024
Show all 44 references
-
[9]
Z. Liu, Y. Zheng, W. Luo. (2025). ArithmeticGPT: empowering small-size large language models with advanced arithmetic skills.Machine Learning,114(1), 2025, 24
2025
-
[10]
J. Gao, R. Pi, J. Zhang, J. Ye, W. Zhong, Y. Wang, L. Hong, J. Han, H. Xu, Z. Li, and L. Kong. G-LLaVa: Solving geometric problem with Multi-Modal Large Language Model.arXiv preprint, arXiv:2312.11370, 2023
2023 arXiv
-
[11]
Shindo, V
H. Shindo, V. Pfanschilling, K. Kersting. Learning differentiable logic programs for abstract visual reasoning.Machine Learning,113(11–12), 2024, 8533–8584
2024
-
[12]
Kangaroo Mathematics Competition, 2025
Kangaroo Sans Frontieres. Kangaroo Mathematics Competition, 2025
2025
-
[13]
K. Yan, Y. Xu, Z. Du, X. Yao, Z. Wang, X. Guo, and J. Chen. Recitation over reasoning: How cutting-edge language models can fail on elementary school-level reasoning problems? 2025
2025
-
[14]
Rhomrasi, Y
L. Rhomrasi, Y. Ahsini, A. Igualde-S´ aez, and R. et al Vinuesa. LLM performance on mathematical reasoning in Catalan language.Results Eng., page 104366, 2025
2025
-
[15]
M. Seo, H. Hajishirzi, A. Farhadi, O. Etzioni, and C. Malcolm. Solving geometry problems: Combining text and diagram interpretation. InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1466–1476, 2015
2015
-
[16]
Sachan and E
M. Sachan and E. Xing. Learning to solve geometry problems from natural language demonstrations in textbooks. InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (*SEM), pages 251–261, 2017
2017
-
[17]
Alayrac, J
J.B. Alayrac, J. Donahue, P. Luc, A. Miech, and et al. Flamingo: a visual language model for few-shot learning. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 35, pages 24245–24259, 2022
2022
-
[18]
Antol, A
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh. VQA: Visual Question Answering. InProceedings of the IEEE International Conference on Computer Vision (ICCV), pages 2425–2433, 2015
2015
-
[19]
P. Lu, R. Gong, S. Jiang, L. Qiu, S. Huang, X. Liang, and S.-C. Zhu. Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning.arXiv preprint arXiv:2105.04165, 2021
2021 arXiv
-
[20]
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023
-
[21]
Zhang, D
R. Zhang, D. Jiang, Y. Zhang, H. Lin, and et al. Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186. Springer, 2024
2024
-
[22]
L. Fan, W. Hua, X. Li, K. Zhu, M. Jin, L. Li, H. Ling, J. Chi, J. Wang, X. Ma, et al. Evaluating Visual Mathematics Problems with MLLMs16 Nphardeval4v: A dynamic reasoning benchmark of multimodal large language models.arXiv preprint arXiv:2403.01777, 2024
2024 arXiv
-
[23]
Z. Zhou, S. Liu, M. Ning, W. Liu, and et al. Is your model really a good math reasoner? evaluating mathematical reasoning with checklist.arXiv preprint arXiv:2407.08733, 2024
2024 arXiv
-
[24]
Goyal, T
Y. Goyal, T. Khot, D.s Summers-Stay, D. Batra, and D. Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. InProc. IEEE Conf. Comput. Vis. Pattern Recognit., pages 6904–6913, 2017
2017
-
[25]
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Doll´ ar, and C.L. Zitnick. Microsoft COCO captions: Data collection and evaluation server.arXiv preprint arXiv:1504.00325, 2015
2015 arXiv
-
[26]
Saikh, T
T. Saikh, T. Ghosal, A. Mittal, A. Ekbal, and P. Bhattacharyya. ScienceQA: A novel resource for question answering on scholarly articles.Int. J. Digit. Libr., 23(3):289–301, 2022
2022
-
[27]
Caffagni, F
D. Caffagni, F. Cocchi, L. Barsellotti, N. Moratelli, S. Sarto, L. Baraldi, M. Cornia, and R. Cucchiara. The revolution of multimodal large language models: A survey.Findings Assoc. Comput. Linguist., pages 13590–13618, 2024
2024
-
[28]
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, pages 2507–2521, 2022
2022
-
[29]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.arXiv preprint, arXiv:2403.05530, 2024
Google Gemini Team. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.arXiv preprint, arXiv:2403.05530, 2024
2024 arXiv
-
[30]
T. H. Trinh, Y. Wu, Q. V. Le, H. He, and T. Luong. Solving olympiad geometry without human demonstrations.Nature, 625(7995):476–482, 2024
2024
-
[31]
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou. Qwen-VL: A versatile vision-language model for understanding, localization, text reading, and beyond.arXiv preprint, arXiv:2308.12966, 2023
2023 arXiv
-
[32]
Pixtral 12b.Mistral AI News, 2024.https://mistral.ai/news/pixtral-12bLast retrieved, April 25th, 2025
Mistral AI. Pixtral 12b.Mistral AI News, 2024.https://mistral.ai/news/pixtral-12bLast retrieved, April 25th, 2025
2024
-
[33]
Llama 3.2 vision 11b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025
Meta AI. Llama 3.2 vision 11b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025
2024
-
[34]
Llama 3.2 vision 90b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025
Meta AI. Llama 3.2 vision 90b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025
2024
-
[35]
Pixtral large.Mistral AI News, 2024.https://mistral.ai/news/pixtral-large Last retrieved, April 25th, 2025
Mistral AI. Pixtral large.Mistral AI News, 2024.https://mistral.ai/news/pixtral-large Last retrieved, April 25th, 2025
2024
-
[36]
Gpt-4o.OpenAI Documentation, 2024.https://platform.openai.com/docs/models Last retrieved, April 25th, 2025
OpenAI. Gpt-4o.OpenAI Documentation, 2024.https://platform.openai.com/docs/models Last retrieved, April 25th, 2025
2024
-
[37]
Gemini 2.0 flash.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash/Last retrieved, April 25th, 2025
Google AI. Gemini 2.0 flash.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash/Last retrieved, April 25th, 2025
2024
-
[38]
Gemini 2.0 flash lite.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash-lite/Last retrieved, April 25th, 2025
Google AI. Gemini 2.0 flash lite.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash-lite/Last retrieved, April 25th, 2025
2024
-
[39]
Qwen2.5-vl 72b.GitHub Repository, 2024.https://github.com/QwenLM/Qwen2.5-VL Last retrieved, April 25th, 2025
QwenLM. Qwen2.5-vl 72b.GitHub Repository, 2024.https://github.com/QwenLM/Qwen2.5-VL Last retrieved, April 25th, 2025
2024
-
[40]
Australian Math Trust, 2025.https://www.amt.edu.au/ Last retrieved, April 25th, 2025
Australian Mathematics Competition. Australian Math Trust, 2025.https://www.amt.edu.au/ Last retrieved, April 25th, 2025
2025
-
[41]
Kangaroo mathematics competition, 2025
European Mathematical Society and Kangourou Sans Fronti` ere. Kangaroo mathematics competition, 2025
2025
-
[42]
Concurso Canguro de Matem´ aticas, 2025.https://canguromat.es/Last retrieved, April 25th, 2025
Federaci´ on Espa˜ nola de Sociedades de Profesores de Matem´ aticas. Concurso Canguro de Matem´ aticas, 2025.https://canguromat.es/Last retrieved, April 25th, 2025
2025
-
[43]
Concours Kangourou de Math´ ematiques, 2025.https://www.aksf
Kangourou Sans Fronti` ere. Concours Kangourou de Math´ ematiques, 2025.https://www.aksf. org/Last retrieved, April 25th, 2025
2025
-
[44]
Concurs Cangur de Matem` atiques, 2025.https://scm.iec
Societat Catalana de Matem` atiques. Concurs Cangur de Matem` atiques, 2025.https://scm.iec. cat/concurs-presencial/cangur/Last retrieved, April 25th, 2025
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.