Agents & computer use
GDPval-AA v2
Artificial Analysis's agentic evaluation of OpenAI's GDPval task set, covering real-world knowledge work across dozens of occupations and industries. Scored via blind pairwise comparison fit to a Bradley-Terry Elo model anchored at 1000 for human-expert deliverables; requires shell and browser access through an agent harness (Stirrup), so it is not a raw model-only capability score.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Agents & computer use
- Metric
- Elo
- Direction
- Higher is better
- Coverage
- 24 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1Claude Fable 5.11,853
Anthropicclosed
- 2Claude Opus 51,824
Anthropicclosed
- 3GLM-5.3 Flash1,765
Zhipu AIopen
- 4GLM-5.31,758
Zhipu AIopen
- 5Grok 4.61,755
xAIclosed
- 6Muse Spark 1.31,754
Metaclosed
- 7Qwen3.8 Flash Next1,743
Alibabaopen
- 8Qwen3.8 Max1,721
Alibabaopen
- 9GPT-5.6 Sol1,710
OpenAIclosed
- 10Kimi K31,668
Moonshotopen
- 11Muse Spark 1.21,615
Metaclosed
- 12Claude Sonnet 51,584
Anthropicclosed
- 13Claude Opus 4.81,578
Anthropicclosed
- 14DeepSeek V4 Pro 08131,577
DeepSeekclosed
- 15GPT-5.6 Luna1,569
OpenAIclosed
- 16GPT-5.6 Terra1,566
OpenAIclosed
- 17DeepSeek V4 Flash 07311,547
DeepSeekopen
- 18Gemini 3.8 Flash1,545
Googleclosed
- 19Qwen3.8 27B1,543
Alibabaopen
- 20Grok 4.51,518
xAIclosed
- 21Gemini 3.7 Flash1,516
Googleclosed
- 22GLM-5.21,498
Zhipu AIopen
- 23Claude Opus 4.71,483
Anthropicclosed
- 24GPT-5.51,482
OpenAIclosed