LLMcompare

Search

Search for a command to run...

Agents & computer use

AutomationBench (Public)

Zapier benchmark evaluating how reliably agents complete realistic business workflows across 47 simulated SaaS tools; scores are agent-system results and depend on the scaffold.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Agents & computer use
Metric
Percent
Direction
Higher is better
Coverage
16 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Grok 4.7

    xAIclosed

    66.0%
  2. 2
    Atria Dawn Preview

    Shanghai AI Labopen

    53.8%
  3. 3
    GLM-5.3

    Zhipu AIopen

    48.2%
  4. 4
    GPT-6 Astra

    OpenAIclosed

    41.4%
  5. 5
    Claude Opus 5.5

    Anthropicclosed

    40.0%
  6. 6
    GPT-6.1 Sol

    OpenAIclosed

    36.1%
  7. 7
    GPT-6 Sol

    OpenAIclosed

    33.2%
  8. 8
    DeepSeek V4 Pro 0813

    DeepSeekclosed

    31.8%
  9. 9
    Claude Fable 5.1

    Anthropicclosed

    31.4%
  10. 10
    Qwen3.8 Max

    Alibabaopen

    27.3%
  11. 11
    Claude Opus 5

    Anthropicclosed

    26.9%
  12. 1225.1%
  13. 13
    GPT-5.6 Sol

    OpenAIclosed

    18.1%
  14. 14
    Claude Fable 5

    Anthropicclosed

    17.4%
  15. 15
    DeepSeek V4 Pro

    DeepSeekopen

    12.8%
  16. 16
    DeepSeek V4 Flash

    DeepSeekopen

    10.8%

Back to all benchmarks