BLUEX v2 creates a public dataset from 2022-2025 UNICAMP and USP second-phase exams and evaluates 21 LLMs via LLM-as-judge, finding scores from 4.18 to 9.10 with math and image understanding as weakest areas.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams
BLUEX v2 creates a public dataset from 2022-2025 UNICAMP and USP second-phase exams and evaluates 21 LLMs via LLM-as-judge, finding scores from 4.18 to 9.10 with math and image understanding as weakest areas.