LLMcompare

Search

Search for a command to run...

Agents & computer use

Terminal-Bench-Science 0.1

Agentic scientific-research workflows across the life, physical, Earth, mathematical, and engineering sciences, built on the Terminal-Bench harness with tasks authored by domain experts and verified programmatically. Scores are agent-system results with a wide per-model standard error (roughly ±3.5-4.5 pts) and are not comparable across benchmark releases.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Agents & computer use
Metric
Percent
Direction
Higher is better
Coverage
15 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    GPT-6 Astra

    OpenAIclosed

    64.6%
  2. 2
    Claude Opus 5.5

    Anthropicclosed

    58.7%
  3. 3
    GPT-6.1 Sol

    OpenAIclosed

    57.0%
  4. 4
    Claude Fable 5.1

    Anthropicclosed

    52.6%
  5. 5
    Claude Opus 5

    Anthropicclosed

    30.0%
  6. 6
    GPT-5.6 Sol

    OpenAIclosed

    22.4%
  7. 7
    Claude Fable 5

    Anthropicclosed

    21.4%
  8. 8
    Gemini 3.8 Flash

    Googleclosed

    12.4%
  9. 9
    Claude Opus 4.8

    Anthropicclosed

    10.5%
  10. 10
    GPT-5.6 Terra

    OpenAIclosed

    8.6%
  11. 11
    GLM-5.3

    Zhipu AIopen

    8.1%
  12. 12
    Grok 4.6

    xAIclosed

    7.1%
  13. 13
    Kimi K3

    Moonshotopen

    7.1%
  14. 14
    Gemini 3.7 Flash

    Googleclosed

    5.7%
  15. 15
    GPT-5.6 Luna

    OpenAIclosed

    3.3%

Back to all benchmarks