BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams
For researchers and practitioners evaluating LLMs in Portuguese, this benchmark fills a gap by providing open-ended, multimodal questions that test deeper reasoning beyond multiple-choice.
BLUEX v2 introduces a benchmark of 395 open-ended questions from Brazilian university entrance exams, with 919 graded subquestions, to evaluate LLMs on discursive reasoning in Portuguese. Evaluation of 21 models shows a 4.92-point performance spread (4.18-9.10 out of 10), with Mathematical Reasoning and Image Understanding being the hardest capabilities.
Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, discursive tasks that demand deeper reasoning and generation capabilities. While the original BLUEX benchmark addressed the scarcity of Portuguese evaluation datasets through multiple-choice questions from Brazilian university entrance exams, it did not cover the more challenging second-phase examinations, which require free-form written responses. In this work, we introduce BLUEX v2, a benchmark derived from the second-phase entrance exams of Brazil's two leading universities: UNICAMP (Comvest) and USP (Fuvest), spanning exam years 2022-2025. Our dataset comprises 395 questions unfolding into 919 graded subquestions, with 55.7% of questions containing associated images. Each question is annotated with subject area, official reference answers, LLM-generated rubric criteria, and six cognitive capability tags. We evaluate 21 state-of-the-art LLMs using an LLM-as-a-judge protocol. Results reveal a 4.92-point performance spread across models (4.18-9.10 on a 0-10 scale), with Mathematical Reasoning and Image Understanding emerging as the hardest capability dimensions. The dataset, evaluation code, and model outputs are publicly available at https://anonymous.4open.science/r/BLUEXv2.