Pith. sign in

REVIEW 6 cited by

OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12753 v2 pith:MWIGUTQR submitted 2024-06-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoningcognitivemodelschallengescomplexabilitiesacrossolympicarena
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abilities in problem-solving and scientific discovery (i.e., AI4Science) once exclusive to human intellect. To comprehensively evaluate current models' performance in cognitive reasoning abilities, we introduce OlympicArena, which includes 11,163 bilingual problems across both text-only and interleaved text-image modalities. These challenges encompass a wide range of disciplines spanning seven fields and 62 international Olympic competitions, rigorously examined for data leakage. We argue that the challenges in Olympic competition problems are ideal for evaluating AI's cognitive reasoning due to their complexity and interdisciplinary nature, which are essential for tackling complex scientific challenges and facilitating discoveries. Beyond evaluating performance across various disciplines using answer-only criteria, we conduct detailed experiments and analyses from multiple perspectives. We delve into the models' cognitive reasoning abilities, their performance across different modalities, and their outcomes in process-level evaluations, which are vital for tasks requiring complex reasoning with lengthy solutions. Our extensive evaluations reveal that even advanced models like GPT-4o only achieve a 39.97% overall accuracy, illustrating current AI limitations in complex reasoning and multimodal integration. Through the OlympicArena, we aim to advance AI towards superintelligence, equipping it to address more complex challenges in science and beyond. We also provide a comprehensive set of resources to support AI research, including a benchmark dataset, an open-source annotation platform, a detailed evaluation tool, and a leaderboard with automatic submission features.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A two-track Olympiad benchmark with deterministic integer answers and step-by-step proof grading shows frontier LLMs drop sharply versus older math benchmarks.

  2. Multilingual Multimodal Software Developer for Code Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 7B vision-language model trained on synthetic diagram-to-code data outperforms several larger open-weight models on a new 10-language UML/flowchart code-generation benchmark.

  3. ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark of perturbed symbolic math problems shows that large language models' performance drops sharply under minor numeric, symbolic, and equivalence transformations.

  4. UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    UGPhysics is a new bilingual benchmark of 5,520 undergraduate physics problems; the strongest tested LLM, OpenAI o1-mini, reaches only 49.8% accuracy.

  5. SciDA: Scientific Dynamic Assessor of LLMs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    SciDA is a dynamically initialized, multi-discipline olympiad benchmark that shows LLMs perform substantially worse when problem variables are randomized, which the authors attribute to memorization of fixed numerical...

  6. Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A thesis that combines self-learning from dialog logs, schema-guided prompting, and self-aligned factuality to build task bots with minimal human intervention.

Pith tools