LLMcompare

Search

Search for a command to run...

Coding

CursorBench

Cursor's IDE-native agent eval on ambiguous, multi-file tasks from real Cursor sessions — correctness under Cursor's agent scaffolding (v3.2), not a model-only test.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
20 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Fable 5.1

    Anthropicclosed

    73.4%
  2. 2
    Grok 4.6

    xAIclosed

    70.8%
  3. 3
    Claude Fable 5

    Anthropicclosed

    70.5%
  4. 4
    Claude Mythos 5

    Anthropicclosed

    70.5%
  5. 5
    Claude Opus 5

    Anthropicclosed

    70.0%
  6. 6
    Gemini 3.8 Flash

    Googleclosed

    69.2%
  7. 7
    GPT-5.6 Sol

    OpenAIclosed

    67.2%
  8. 8
    Cursor Grok 4.5

    Cursorclosed

    66.7%
  9. 9
    GPT-5.6 Terra

    OpenAIclosed

    64.9%
  10. 10
    Claude Opus 4.8

    Anthropicclosed

    62.3%
  11. 11
    Gemini 3.7 Flash

    Googleclosed

    61.6%
  12. 12
    Claude Sonnet 5

    Anthropicclosed

    61.5%
  13. 13
    GPT-5.6 Luna

    OpenAIclosed

    61.1%
  14. 14
    Kimi K3

    Moonshotopen

    60.8%
  15. 15
    GPT-5.5

    OpenAIclosed

    58.4%
  16. 16
    Composer 2.5

    Cursorclosed

    56.1%
  17. 17
    GLM-5.2

    Zhipu AIopen

    55.0%
  18. 18
    Gemini 3.6 Flash

    Googleclosed

    53.5%
  19. 19
    Kimi K2.7 Code

    Moonshotopen

    49.7%
  20. 20
    Gemini 3.5 Flash

    Googleclosed

    48.8%

Back to all benchmarks