Agents & computer use
OSWorld 2.0 (Strict)
OSWorld 2.0 long-horizon computer-use benchmark under strict scoring, which requires the full task to complete for credit. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to partial-credit OSWorld 2.0 scores or to pre-2.0 OSWorld results.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Agents & computer use
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 3 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
- Source
- os-world.github.io/
Scores
- 1Claude Fable 5.141.7%
Anthropicclosed
- 2Claude Opus 539.6%
Anthropicclosed
- 3Claude Fable 536.1%
Anthropicclosed