Agents & computer use
Terminal-Bench-Science 0.1
Agentic scientific-research workflows across the life, physical, Earth, mathematical, and engineering sciences, built on the Terminal-Bench harness with tasks authored by domain experts and verified programmatically. Scores are agent-system results with a wide per-model standard error (roughly ±3.5-4.5 pts) and are not comparable across benchmark releases.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Agents & computer use
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 15 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1GPT-6 Astra64.6%
OpenAIclosed
- 2Claude Opus 5.558.7%
Anthropicclosed
- 3GPT-6.1 Sol57.0%
OpenAIclosed
- 4Claude Fable 5.152.6%
Anthropicclosed
- 5Claude Opus 530.0%
Anthropicclosed
- 6GPT-5.6 Sol22.4%
OpenAIclosed
- 7Claude Fable 521.4%
Anthropicclosed
- 8Gemini 3.8 Flash12.4%
Googleclosed
- 9Claude Opus 4.810.5%
Anthropicclosed
- 10GPT-5.6 Terra8.6%
OpenAIclosed
- 11GLM-5.38.1%
Zhipu AIopen
- 12Grok 4.67.1%
xAIclosed
- 13Kimi K37.1%
Moonshotopen
- 14Gemini 3.7 Flash5.7%
Googleclosed
- 15GPT-5.6 Luna3.3%
OpenAIclosed