LLMcompare

Search

Search for a command to run...

Math

MathArena

MathArena contest-math leaderboard (AIME, AMC, olympiad-style); results depend on answer-format extraction and evaluation date.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Math
Metric
Percent
Direction
Higher is better
Coverage
4 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Opus 5

    Anthropicclosed

    84.4%
  2. 2
    GPT-5.6 Sol

    OpenAIclosed

    79.7%
  3. 3
    GPT-5.5

    OpenAIclosed

    77.8%
  4. 4
    Kimi K3

    Moonshotopen

    69.7%

Back to all benchmarks