LLMcompare

Search

Search for a command to run...

Agents & computer use

OSWorld 2.0 (Strict)

OSWorld 2.0 long-horizon computer-use benchmark under strict scoring, which requires the full task to complete for credit. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to partial-credit OSWorld 2.0 scores or to pre-2.0 OSWorld results.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Agents & computer use
Metric
Percent
Direction
Higher is better
Coverage
3 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Fable 5.1

    Anthropicclosed

    41.7%
  2. 2
    Claude Opus 5

    Anthropicclosed

    39.6%
  3. 3
    Claude Fable 5

    Anthropicclosed

    36.1%

Back to all benchmarks