LLMcompare

Search

Search for a command to run...

Agents & computer use

GDPval-AA v2

Artificial Analysis's agentic evaluation of OpenAI's GDPval task set, covering real-world knowledge work across dozens of occupations and industries. Scored via blind pairwise comparison fit to a Bradley-Terry Elo model anchored at 1000 for human-expert deliverables; requires shell and browser access through an agent harness (Stirrup), so it is not a raw model-only capability score.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Agents & computer use
Metric
Elo
Direction
Higher is better
Coverage
24 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Fable 5.1

    Anthropicclosed

    1,853
  2. 2
    Claude Opus 5

    Anthropicclosed

    1,824
  3. 3
    GLM-5.3 Flash

    Zhipu AIopen

    1,765
  4. 4
    GLM-5.3

    Zhipu AIopen

    1,758
  5. 5
    Grok 4.6

    xAIclosed

    1,755
  6. 6
    Muse Spark 1.3

    Metaclosed

    1,754
  7. 71,743
  8. 8
    Qwen3.8 Max

    Alibabaopen

    1,721
  9. 9
    GPT-5.6 Sol

    OpenAIclosed

    1,710
  10. 10
    Kimi K3

    Moonshotopen

    1,668
  11. 11
    Muse Spark 1.2

    Metaclosed

    1,615
  12. 12
    Claude Sonnet 5

    Anthropicclosed

    1,584
  13. 13
    Claude Opus 4.8

    Anthropicclosed

    1,578
  14. 14
    DeepSeek V4 Pro 0813

    DeepSeekclosed

    1,577
  15. 15
    GPT-5.6 Luna

    OpenAIclosed

    1,569
  16. 16
    GPT-5.6 Terra

    OpenAIclosed

    1,566
  17. 171,547
  18. 18
    Gemini 3.8 Flash

    Googleclosed

    1,545
  19. 19
    Qwen3.8 27B

    Alibabaopen

    1,543
  20. 20
    Grok 4.5

    xAIclosed

    1,518
  21. 21
    Gemini 3.7 Flash

    Googleclosed

    1,516
  22. 22
    GLM-5.2

    Zhipu AIopen

    1,498
  23. 23
    Claude Opus 4.7

    Anthropicclosed

    1,483
  24. 24
    GPT-5.5

    OpenAIclosed

    1,482

Back to all benchmarks