Agents & computer use
AutomationBench (Public)
Zapier benchmark evaluating how reliably agents complete realistic business workflows across 47 simulated SaaS tools; scores are agent-system results and depend on the scaffold.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Agents & computer use
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 16 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1Grok 4.766.0%
xAIclosed
- 2Atria Dawn Preview53.8%
Shanghai AI Labopen
- 3GLM-5.348.2%
Zhipu AIopen
- 4GPT-6 Astra41.4%
OpenAIclosed
- 5Claude Opus 5.540.0%
Anthropicclosed
- 6GPT-6.1 Sol36.1%
OpenAIclosed
- 7GPT-6 Sol33.2%
OpenAIclosed
- 8DeepSeek V4 Pro 081331.8%
DeepSeekclosed
- 9Claude Fable 5.131.4%
Anthropicclosed
- 10Qwen3.8 Max27.3%
Alibabaopen
- 11Claude Opus 526.9%
Anthropicclosed
- 12DeepSeek V4 Flash 073125.1%
DeepSeekopen
- 13GPT-5.6 Sol18.1%
OpenAIclosed
- 14Claude Fable 517.4%
Anthropicclosed
- 15DeepSeek V4 Pro12.8%
DeepSeekopen
- 16DeepSeek V4 Flash10.8%
DeepSeekopen