Coding
Terminal-Bench 3
Next-generation agentic terminal benchmark (Frontier-Bench); scores are agent-system results that vary with harness, effort, and evaluation date.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Coding
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 13 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
- Source
- www.frontierbench.ai/
Scores
- 1Claude Opus 542.7%
Anthropicclosed
- 2GPT-5.6 Sol34.6%
OpenAIclosed
- 3Claude Fable 534.1%
Anthropicclosed
- 4Claude Mythos 534.1%
Anthropicclosed
- 5GLM-5.328.3%
Zhipu AIopen
- 6Grok 4.626.5%
xAIclosed
- 7Claude Opus 4.821.1%
Anthropicclosed
- 8GPT-5.6 Terra20.8%
OpenAIclosed
- 9Grok 4.515.7%
xAIclosed
- 10Gemini 3.7 Flash14.9%
Googleclosed
- 11Claude Sonnet 514.6%
Anthropicclosed
- 12GPT-5.6 Luna14.3%
OpenAIclosed
- 13GLM-5.24.6%
Zhipu AIopen