LLMcompare

Search

Search for a command to run...

Agents & computer use

OSWorld 2.0 (Partial Credit)

OSWorld 2.0 long-horizon computer-use benchmark under partial-credit scoring, which awards credit for subtasks completed within a long real-world desktop workflow. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to pre-2.0 OSWorld results.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Agents & computer use
Metric
Percent
Direction
Higher is better
Coverage
9 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Opus 5.5

    Anthropicclosed

    81.8%
  2. 2
    Claude Fable 5.1

    Anthropicclosed

    77.9%
  3. 3
    Claude Fable 5

    Anthropicclosed

    72.9%
  4. 4
    GPT-6 Astra

    OpenAIclosed

    72.6%
  5. 5
    GPT-6.1 Sol

    OpenAIclosed

    71.4%
  6. 6
    Claude Opus 5

    Anthropicclosed

    70.2%
  7. 7
    GPT-5.6 Sol

    OpenAIclosed

    65.7%
  8. 8
    GPT-6 Sol

    OpenAIclosed

    60.5%
  9. 9
    Gemini 3.8 Flash

    Googleclosed

    59.0%

Back to all benchmarks