Math
FrontierMath Tier 4 (v2)
Epoch AI's FrontierMath benchmark, Tier 4 (v2 problem set): exceptionally difficult, original research-level mathematics problems vetted by expert mathematicians. Tier 4 is not comparable to FrontierMath Tiers 1-3, and v2 problem sets are not comparable to v1 results.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Math
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 5 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1GPT-6 Astra97.6%
OpenAIclosed
- 2Claude Fable 5.187.8%
Anthropicclosed
- 3Claude Fable 587.8%
Anthropicclosed
- 4GPT-5.6 Sol83.0%
OpenAIclosed
- 5Claude Opus 573.2%
Anthropicclosed