Agents & computer use
Agents' Last Exam
Agent benchmark of long-horizon, real-world tasks requiring tool use and multi-step execution; results are agent-system measurements published by the evaluating provider.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Agents & computer use
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 10 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
- Source
- agents-last-exam.org/
Scores
- 1GPT-6 Astra59.3%
OpenAIclosed
- 2GPT-6 Sol56.4%
OpenAIclosed
- 3Claude Opus 555.5%
Anthropicclosed
- 4GPT-5.6 Sol53.6%
OpenAIclosed
- 5Claude Fable 548.7%
Anthropicclosed
- 6GLM-5.328.5%
Zhipu AIopen
- 7DeepSeek V4 Pro 081325.7%
DeepSeekclosed
- 8DeepSeek V4 Flash 073125.2%
DeepSeekopen
- 9DeepSeek V4 Pro16.5%
DeepSeekopen
- 10DeepSeek V4 Flash15.8%
DeepSeekopen