Pith. sign in

Paper Citation Record · LEDGER

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests

As of 14 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2506.07418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07418 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:39:00.747897Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a484e773-47e5-42ae-b5d5-8fb790abd933 · outbound

This paper cites Zhang, Y.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Zhang, Y

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:06.503087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:56.536458Z digest=sha256:dd6aa0d16e68a806e2042566fe3801d2ad6e9badf2687017cd1027898af6ac38

Observation 3aae2431-7507-489e-bea7-b08a9baa8cb8 · outbound

This paper cites Achiam, S.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Achiam, S

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:06.313361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:56.650734Z digest=sha256:e01814233146fc5ff22470759c55da357c51eaf49dda001a8a48b594e2ee09ae

Observation fc18f617-188d-4cea-a35f-ca373b0d3f75 · outbound

This paper cites Qwen2.5-vl 7b.GitHub Repository, 2024.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Qwen2.5-vl 7b.GitHub Repository, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:06.069406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:56.737402Z digest=sha256:5a00924bf3f890369c24acd486d68a38723c87e1a1928104829578fae8ebde64

Observation 42929a52-a881-4b7b-bc17-e02c0c02a3e4 · outbound

This paper cites Visual Instruction Tuning.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Visual Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:56.825172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:56.825172Z digest=sha256:7ebae5677ff89d4d761e37227246bf0d1bce1c73ebd571d7533ab928cdca1f11

Observation 2630dfda-3d82-47c3-8b56-eef2f69ea43e · outbound

This paper cites Andreescu and R.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Andreescu and R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:05.898084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:56.904032Z digest=sha256:bb1530eda2071088662662cfea4a4c61ac5fd435e563770061f55c636e83c12a

Observation 93700d1e-7cf0-4a5d-8199-fbe04144cfe1 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests A Survey on Benchmarks of Multimodal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.018980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.018980Z digest=sha256:821d2186521f30f8241d2c43b568bff3fa65d1e0e82f73ac866d85d323305ff7

Observation 50ca7b15-dfdc-46db-8798-0ddf1a0e0329 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.150799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.150799Z digest=sha256:53a5e467e4cd48cddc03993f1df95c2734a01a4b165b77f91c61987a399cff9a

Observation 87b1b2c3-68b9-4d65-afbc-738f1b3f6947 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.238958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.238958Z digest=sha256:a8953b39d4f8e62ba792bd08da8ca78772fd5d103e19225b651da49edaac6529

Observation cad49236-1d9d-49f3-9a3f-c38b95ef13a2 · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:05.722951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:57.326879Z digest=sha256:d9278dde7035bd440b5f393ba30b260943990e153ee18e77f959ac3aa48c196e

Observation 304fed1b-577f-4a56-9216-74a69416e9ca · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.399389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.399389Z digest=sha256:1b8b2b6780de6e3f3358c012e6f715e1c41aec33f41b356db629c117ad438c5a

Observation 5138c6dd-79fd-47d8-89ba-261db7c90624 · outbound

This paper cites Shindo, V.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Shindo, V

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:05.485076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:57.478028Z digest=sha256:fe83d39fce4e652034b34c2d0265053666d32082f2a678c2041ddb12ef732a7e

Observation caf37e07-1b82-4459-a02c-4927935021c0 · outbound

This paper cites Kangaroo Mathematics Competition, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Kangaroo Mathematics Competition, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:05.273555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:57.585502Z digest=sha256:f9b7972c96e58e5c3523a6a54d8be150eb189c468d8edd78c4a3e5cd799a1418

Observation cb88a8b8-03b8-4266-8d7e-738e2c7f8f5b · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:05.148045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:57.678482Z digest=sha256:5684abfe5dadf7604b953224c17963821fa05464aaf583bd5819e1322032791c

Observation 573b588d-a2f9-42f0-9b3f-f827c66c52d0 · outbound

This paper cites Rhomrasi, Y.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Rhomrasi, Y

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.938401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:57.785035Z digest=sha256:1b79a728a55e5610826351c1d872b5353b306e37fc9b9153a4142f5acfb4e872

Observation d745f0ea-b10e-4459-9681-3fd338604861 · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:04.752925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:57.900519Z digest=sha256:e3aa0c2836a7179fe597861c8ff16df6def455687e8ed5812c218777590754d9

Observation 2d0e39b0-00e1-4238-ba58-8e2830ac619f · outbound

This paper cites Sachan and E.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Sachan and E

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.548848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:58.004348Z digest=sha256:3a919d0e788765c46d1b5e33ba0d367bef5a1b1d480cb506903f6753a5aabb5e

Observation c3b793c7-6a0e-4711-ade6-818b4fb0e3c2 · outbound

This paper cites Alayrac, J.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Alayrac, J

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.364881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:58.123777Z digest=sha256:d094dda7946cce184969f193061389e38c96e354531ed88ee92632c54cd8dd76

Observation 66b9ed53-75f3-4785-ad3e-e3e9985d2472 · outbound

This paper cites Antol, A.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Antol, A

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.200993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:58.228933Z digest=sha256:84edf06c36323aa544a6b8373b64469f0c58d5b058a0eccd2a8158d5f21611bd

Observation 1b2bac11-e557-41de-90d9-0663e5af949e · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.315986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.315986Z digest=sha256:325ff824ee8eee529992af4d9ddabaf610802f4527ab6d6c891dfbf58d1b788a

Observation c7c7c834-d8c3-4486-a128-551a39b058d1 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.407354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.407354Z digest=sha256:b658d9f331d073d1cd8664cb6001f133ca0f9822d04ee3ae759072a8057360af

Observation 95d62a69-1cec-46fb-91d8-974add5e83dc · outbound

This paper cites Zhang, D.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Zhang, D

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.016375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:58.556096Z digest=sha256:eeaf15295104c45abfefac05c39b164f7b064e3cf6ceaf1b54b481fee2ae6eaf

Observation 97d54cd4-ff90-40b7-93e9-00ce9a5c2f64 · outbound

This paper cites NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.647891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.647891Z digest=sha256:09abfd03cccbfe56a695da0556265b364b35649388e955545feb5be49ab2a26a

Observation af557481-fa63-43f1-bd55-bc3fbec83de0 · outbound

This paper cites Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.780594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.780594Z digest=sha256:1db7f809737ffe546a39253b7f8627b45424c75cfb7bf9326921aca7fd9e8521

Observation 5f2f6cf7-204e-4d06-8cb0-e32ee99c223a · outbound

This paper cites Goyal, T.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Goyal, T

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.831158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:58.862704Z digest=sha256:f60241f5c476d1e148ed41f332007fb2dd6d7bca3777f3881de32ba2e20a0fda

Observation 453f1ba5-a98e-4c99-b932-26767347630a · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.990977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.990977Z digest=sha256:6d9d0b38440e70ba82028b2df39c479e18f33a969b7583dc96733c0231fcac08

Observation ee210642-10fa-4f89-8037-d4ece739fa93 · outbound

This paper cites Saikh, T.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Saikh, T

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.637710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:59.074449Z digest=sha256:23beeda0638bb2a261a7f7273c8fb9ada9a8bac3c164a1ee3034666cd067f900

Observation 0c87dee2-7f1b-42df-a252-7f049e8e7e4f · outbound

This paper cites Caffagni, F.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Caffagni, F

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.476997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:59.168749Z digest=sha256:b6a9b9f2be2a49aaf7fdbcaa8e5aada59954142388fc73c8f916e054b2688d51

Observation 896556e9-dafe-4f9a-a901-4e5005ad5654 · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:03.348832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:59.258951Z digest=sha256:f2d562cd15ca07194f3bb9998d9333923733fe5f11028576c9cd9536b6a6228b

Observation 5f9d92a7-117f-4199-b4b3-dfe6dd2d72e9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:59.377109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:59.377109Z digest=sha256:f66b98f6d16d87e210e1348a8de23f265288cf6be36b9cd2a78e3eb20927141e

Observation 83ba023d-cbfb-4439-bba7-77341301036e · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:59.464859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:59.464859Z digest=sha256:e9f73abf2574f1515dd5856d1c4a7a951bf4555e36a2251e30db24970d9974e5

Observation 0c8dc9d2-db98-46ab-b742-339bc071c773 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:59.546002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:59.546002Z digest=sha256:24e26cb37b37f79065b2d0ac767a1d079c1af16d1ebf5f47230058813407f977

Observation 690bfc09-49da-4b96-893b-3cf4e768ce37 · outbound

This paper cites Pixtral 12b.Mistral AI News, 2024.https://mistral.ai/news/pixtral-12bLast retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Pixtral 12b.Mistral AI News, 2024.https://mistral.ai/news/pixtral-12bLast retrieved, April 25th, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.227171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:59.623333Z digest=sha256:ae49bac4acb558d1245693d605a104efe7b07fc8bd7605004aa550f4e826da73

Observation 038865bd-3558-44ee-a4bc-43f4220e087e · outbound

This paper cites Llama 3.2 vision 11b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Llama 3.2 vision 11b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.034380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:59.711872Z digest=sha256:801e4bdd0310e3671d9e523ed8f65e16d2d659b337c15ddbfe935276e6298cf4

Observation 6f62abc5-0f3f-46ef-a12e-76aeb0052716 · outbound

This paper cites Llama 3.2 vision 90b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Llama 3.2 vision 90b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.886641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:59.815322Z digest=sha256:500ae2fc6edb58c9a06ae894115d91b76086b92cf448bb5dc6bade7029d63e7d

Observation 3497ccd6-ef16-4153-856e-188d3b84e651 · outbound

This paper cites Pixtral large.Mistral AI News, 2024.https://mistral.ai/news/pixtral-large Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Pixtral large.Mistral AI News, 2024.https://mistral.ai/news/pixtral-large Last retrieved, April 25th, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.746861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:38:59.925316Z digest=sha256:e0bfe178bf426ebb31ed39f59b8fc595f06f8a7c5a39078806d1841657546d15

Observation ab2100e2-8416-4dd2-b4f6-4323ec5a9866 · outbound

This paper cites Gpt-4o.OpenAI Documentation, 2024.https://platform.openai.com/docs/models Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gpt-4o.OpenAI Documentation, 2024.https://platform.openai.com/docs/models Last retrieved, April 25th, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.570734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:39:00.023172Z digest=sha256:cfe59e44b791d1240589ba98dc4f545cfe4e40a8d3ad2057793c365d618246d3

Observation 6a14e7c4-98a9-4cf5-8489-7a1132ca4f4b · outbound

This paper cites Gemini 2.0 flash.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gemini 2.0 flash.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash/Last retrieved, April 25th, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.409137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:39:00.113148Z digest=sha256:bef7fd88136b91746af36ba76bc1233f6749b439c1cf95701185a7a06fc0eae0

Observation 8d7fae9c-9770-4187-81e9-32ad6a13f7b4 · outbound

This paper cites Gemini 2.0 flash lite.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash-lite/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gemini 2.0 flash lite.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash-lite/Last retrieved, April 25th, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.243023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:39:00.238036Z digest=sha256:5ec26b5bffe24d044e26d0a789f7a092802086f9c40606c789f8ffbc894a5f29

Observation f00e5a31-7e5b-41ab-855f-3f97e42af644 · outbound

This paper cites Qwen2.5-vl 72b.GitHub Repository, 2024.https://github.com/QwenLM/Qwen2.5-VL Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Qwen2.5-vl 72b.GitHub Repository, 2024.https://github.com/QwenLM/Qwen2.5-VL Last retrieved, April 25th, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.121010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:39:00.305158Z digest=sha256:68e402c7a54de70c2360768a7f8c5a00cb5aa47b12d0eee444743e9f8d60af47

Observation 1b0fbfde-3c37-4e78-9479-eec4f713b624 · outbound

This paper cites Australian Math Trust, 2025.https://www.amt.edu.au/ Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Australian Math Trust, 2025.https://www.amt.edu.au/ Last retrieved, April 25th, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.924226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:39:00.397520Z digest=sha256:5e01801f1d56e88ee8020e791a3d7d0e982891fcf11f46aa176ac8e850d6a811

Observation d9315d22-26a3-40be-84ad-ba9dff0c6d1d · outbound

This paper cites Kangaroo mathematics competition, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Kangaroo mathematics competition, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.732821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:39:00.468727Z digest=sha256:74ff4090ce1bdb3d52abaacaf05029618e8832b0fbaab25d655502b6bacde35f

Observation cd56db18-9f26-4c6c-8acb-231b9d04268d · outbound

This paper cites Concurso Canguro de Matem´ aticas, 2025.https://canguromat.es/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Concurso Canguro de Matem´ aticas, 2025.https://canguromat.es/Last retrieved, April 25th, 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.559522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:39:00.567258Z digest=sha256:df6dc3e2eb719531af30f64e48d373e7309acdae4252a21db7273eecdfa2223d

Observation 5cd0efb3-23a0-45a0-8ebd-0e81152075a3 · outbound

This paper cites Concours Kangourou de Math´ ematiques, 2025.https://www.aksf.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Concours Kangourou de Math´ ematiques, 2025.https://www.aksf

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.377765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:39:00.656507Z digest=sha256:aaebec5b6d7f2f0dc86fe5f19959ad29622382f64656f77b05a177ce8a84039a

Observation dbb4fb08-ac9f-4904-9217-c03ab09a386f · outbound

This paper cites Concurs Cangur de Matem` atiques, 2025.https://scm.iec.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Concurs Cangur de Matem` atiques, 2025.https://scm.iec

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.164441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:39:00.747897Z digest=sha256:dec16bceb5c716f5922e1be03024abad9f6c0ec776f333e73fe35d9406a2ad3e

Pith citing papers

No inbound Pith citation observations are available.