LLMcompare

Search

Search for a command to run...

Coding

Terminal-Bench 3

Next-generation agentic terminal benchmark (Frontier-Bench); scores are agent-system results that vary with harness, effort, and evaluation date.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
13 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Opus 5

    Anthropicclosed

    42.7%
  2. 2
    GPT-5.6 Sol

    OpenAIclosed

    34.6%
  3. 3
    Claude Fable 5

    Anthropicclosed

    34.1%
  4. 4
    Claude Mythos 5

    Anthropicclosed

    34.1%
  5. 5
    GLM-5.3

    Zhipu AIopen

    28.3%
  6. 6
    Grok 4.6

    xAIclosed

    26.5%
  7. 7
    Claude Opus 4.8

    Anthropicclosed

    21.1%
  8. 8
    GPT-5.6 Terra

    OpenAIclosed

    20.8%
  9. 9
    Grok 4.5

    xAIclosed

    15.7%
  10. 10
    Gemini 3.7 Flash

    Googleclosed

    14.9%
  11. 11
    Claude Sonnet 5

    Anthropicclosed

    14.6%
  12. 12
    GPT-5.6 Luna

    OpenAIclosed

    14.3%
  13. 13
    GLM-5.2

    Zhipu AIopen

    4.6%

Back to all benchmarks