LLMcompare

Search

Search for a command to run...

Coding

Terminal-Bench 4.0

Semantically-versioned continuation of Terminal-Bench with recalibrated task resources and a refreshed, less-saturated task set; scores are agent-system results that vary with harness, effort level, and evaluation date, and are not directly comparable to Terminal-Bench 2.x or 3 results.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
20 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Sonnet 5.5

    Anthropicclosed

    70.6%
  2. 2
    Claude Opus 5.5

    Anthropicclosed

    66.4%
  3. 3
    Claude Mythos 5.1

    Anthropicclosed

    60.9%
  4. 4
    GPT-6 Astra

    OpenAIclosed

    58.2%
  5. 5
    Claude Fable 5.1

    Anthropicclosed

    57.9%
  6. 6
    Claude Opus 5

    Anthropicclosed

    51.8%
  7. 7
    Claude Fable 5

    Anthropicclosed

    44.6%
  8. 8
    GPT-6 Sol

    OpenAIclosed

    44.0%
  9. 9
    GLM-5.3

    Zhipu AIopen

    41.8%
  10. 10
    GPT-5.6 Sol

    OpenAIclosed

    37.3%
  11. 11
    Grok 4.7

    xAIclosed

    26.0%
  12. 12
    Claude Opus 4.8

    Anthropicclosed

    23.6%
  13. 13
    GPT-5.6 Terra

    OpenAIclosed

    21.5%
  14. 14
    Grok 4.6

    xAIclosed

    20.3%
  15. 15
    Gemini 3.8 Flash

    Googleclosed

    19.1%
  16. 16
    GPT-5.6 Luna

    OpenAIclosed

    17.3%
  17. 17
    GPT-6 Luna

    OpenAIclosed

    13.0%
  18. 18
    Claude Sonnet 5

    Anthropicclosed

    12.4%
  19. 19
    Grok 4.5

    xAIclosed

    12.4%
  20. 20
    Gemini 3.7 Flash

    Googleclosed

    11.2%

Back to all benchmarks