LLMcompare

Search

Search for a command to run...

Math

FrontierMath Tier 4 (v2)

Epoch AI's FrontierMath benchmark, Tier 4 (v2 problem set): exceptionally difficult, original research-level mathematics problems vetted by expert mathematicians. Tier 4 is not comparable to FrontierMath Tiers 1-3, and v2 problem sets are not comparable to v1 results.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Math
Metric
Percent
Direction
Higher is better
Coverage
5 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    GPT-6 Astra

    OpenAIclosed

    97.6%
  2. 2
    Claude Fable 5.1

    Anthropicclosed

    87.8%
  3. 3
    Claude Fable 5

    Anthropicclosed

    87.8%
  4. 4
    GPT-5.6 Sol

    OpenAIclosed

    83.0%
  5. 5
    Claude Opus 5

    Anthropicclosed

    73.2%

Back to all benchmarks