LLMcompare

Search

Search for a command to run...

Agents & computer use

AA-Briefcase v1.1

Artificial Analysis's private agentic knowledge-work evaluation across four multi-week projects and 91 tasks. Its combined Elo metric includes rubric pass rate, analytical quality, and presentation quality; scores evaluate the agent system and its harness, not a model alone.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Agents & computer use
Metric
Elo
Direction
Higher is better
Coverage
2 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Sonnet 5.5

    Anthropicclosed

    1,811
  2. 2
    Grok 4.7

    xAIclosed

    1,657

Back to all benchmarks