Composite ranking
Agents & computer use
A single derived ranking across the 4 agents & computer usebenchmarks in this catalog. This is not a published benchmark score — it's computed here.
Methodology
Each eligible benchmark's scores are z-score normalized against models released in the past 12 months (since 2025-09-29), then combined as a shrunken mean: sum(z) / (n + 1), with 1 dummy average results so thin coverage still cannot dominate. Saturated benches and benches scored on fewer than ten peer models are excluded. A model needs a published score on at least 3 of the 4 benchmarks below to receive a composite; every contributing score is shown so nothing is hidden behind the composite number.
- Benchmarks included
- ALE, AutoBench, TB-Science, GDPval-AA
- Min. coverage
- 3 of 4 benchmarks
- Coverage
- 8 modelsModels in this catalog with enough published scores for a composite
- Updated
- Catalog snapshot date
Composite ranking
- ALE: 59.3%
- AutoBench: 41.4%
- TB-Science: 64.6%
- AutoBench: 31.4%
- TB-Science: 52.6%
- GDPval-AA: 1,853
- ALE: 55.5%
- AutoBench: 26.9%
- TB-Science: 30.0%
- GDPval-AA: 1,824
- ALE: 28.5%
- AutoBench: 48.2%
- TB-Science: 8.1%
- GDPval-AA: 1,758
- ALE: 53.6%
- AutoBench: 18.1%
- TB-Science: 22.4%
- GDPval-AA: 1,710
- ALE: 48.7%
- AutoBench: 17.4%
- TB-Science: 21.4%
- ALE: 25.7%
- AutoBench: 31.8%
- GDPval-AA: 1,577
- ALE: 25.2%
- AutoBench: 25.1%
- GDPval-AA: 1,547