LLMcompare

Search

Search for a command to run...

Agents & computer use

Agents' Last Exam

Agent benchmark of long-horizon, real-world tasks requiring tool use and multi-step execution; results are agent-system measurements published by the evaluating provider.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Agents & computer use
Metric
Percent
Direction
Higher is better
Coverage
10 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    GPT-6 Astra

    OpenAIclosed

    59.3%
  2. 2
    GPT-6 Sol

    OpenAIclosed

    56.4%
  3. 3
    Claude Opus 5

    Anthropicclosed

    55.5%
  4. 4
    GPT-5.6 Sol

    OpenAIclosed

    53.6%
  5. 5
    Claude Fable 5

    Anthropicclosed

    48.7%
  6. 6
    GLM-5.3

    Zhipu AIopen

    28.5%
  7. 7
    DeepSeek V4 Pro 0813

    DeepSeekclosed

    25.7%
  8. 825.2%
  9. 9
    DeepSeek V4 Pro

    DeepSeekopen

    16.5%
  10. 10
    DeepSeek V4 Flash

    DeepSeekopen

    15.8%

Back to all benchmarks