Coding
Terminal-Bench 4.0
Semantically-versioned continuation of Terminal-Bench with recalibrated task resources and a refreshed, less-saturated task set; scores are agent-system results that vary with harness, effort level, and evaluation date, and are not directly comparable to Terminal-Bench 2.x or 3 results.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Coding
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 20 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1Claude Sonnet 5.570.6%
Anthropicclosed
- 2Claude Opus 5.566.4%
Anthropicclosed
- 3Claude Mythos 5.160.9%
Anthropicclosed
- 4GPT-6 Astra58.2%
OpenAIclosed
- 5Claude Fable 5.157.9%
Anthropicclosed
- 6Claude Opus 551.8%
Anthropicclosed
- 7Claude Fable 544.6%
Anthropicclosed
- 8GPT-6 Sol44.0%
OpenAIclosed
- 9GLM-5.341.8%
Zhipu AIopen
- 10GPT-5.6 Sol37.3%
OpenAIclosed
- 11Grok 4.726.0%
xAIclosed
- 12Claude Opus 4.823.6%
Anthropicclosed
- 13GPT-5.6 Terra21.5%
OpenAIclosed
- 14Grok 4.620.3%
xAIclosed
- 15Gemini 3.8 Flash19.1%
Googleclosed
- 16GPT-5.6 Luna17.3%
OpenAIclosed
- 17GPT-6 Luna13.0%
OpenAIclosed
- 18Claude Sonnet 512.4%
Anthropicclosed
- 19Grok 4.512.4%
xAIclosed
- 20Gemini 3.7 Flash11.2%
Googleclosed