Agents & computer use
AA-Briefcase v1.1
Artificial Analysis's private agentic knowledge-work evaluation across four multi-week projects and 91 tasks. Its combined Elo metric includes rubric pass rate, analytical quality, and presentation quality; scores evaluate the agent system and its harness, not a model alone.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Agents & computer use
- Metric
- Elo
- Direction
- Higher is better
- Coverage
- 2 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1Claude Sonnet 5.51,811
Anthropicclosed
- 2Grok 4.71,657
xAIclosed