Agents & computer use
OSWorld 2.0 (Partial Credit)
OSWorld 2.0 long-horizon computer-use benchmark under partial-credit scoring, which awards credit for subtasks completed within a long real-world desktop workflow. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to pre-2.0 OSWorld results.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Agents & computer use
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 9 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
- Source
- os-world.github.io/
Scores
- 1Claude Opus 5.581.8%
Anthropicclosed
- 2Claude Fable 5.177.9%
Anthropicclosed
- 3Claude Fable 572.9%
Anthropicclosed
- 4GPT-6 Astra72.6%
OpenAIclosed
- 5GPT-6.1 Sol71.4%
OpenAIclosed
- 6Claude Opus 570.2%
Anthropicclosed
- 7GPT-5.6 Sol65.7%
OpenAIclosed
- 8GPT-6 Sol60.5%
OpenAIclosed
- 9Gemini 3.8 Flash59.0%
Googleclosed