LLMcompare

Search

Search for a command to run...

Coding

SWE-bench Verified

Resolves real GitHub issues in popular Python repositories (500-task human-verified subset); agent scaffold, patch budget, and test policy materially affect the result.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
39 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Mythos 5

    Anthropicclosed

    86.3%
  2. 2
    Gemini 3.1 Pro

    Googleclosed

    80.6%
  3. 3
    Qwen3.7 Max

    Alibabaclosed

    80.4%
  4. 4
    Inkling-Small

    Thinking Machinesopen

    80.2%
  5. 5
    Inkling

    Thinking Machinesopen

    77.6%
  6. 6
    Kimi K2.5

    Moonshotopen

    76.8%
  7. 7
    Muse Glimmer

    Metaopen

    76.0%
  8. 8
    MiniMax M2.5

    MiniMaxopen

    75.8%
  9. 9
    Claude Opus 4.6

    Anthropicclosed

    75.6%
  10. 10
    GPT-5

    OpenAIclosed

    74.9%
  11. 11
    Claude Opus 4.1

    Anthropicclosed

    74.5%
  12. 12
    Claude Haiku 4.5

    Anthropicclosed

    73.3%
  13. 13
    GLM-5

    Zhipu AIopen

    72.8%
  14. 1470.7%
  15. 15
    DeepSeek V3.2

    DeepSeekopen

    70.0%
  16. 16
    Gemini 3 Pro

    Googleclosed

    69.6%
  17. 17
    Kimi K2

    Moonshotopen

    65.8%
  18. 18
    gpt-oss-20B

    OpenAIopen

    60.7%
  19. 19
    o3

    OpenAIclosed

    58.4%
  20. 20
    DeepSeek R1-0528

    DeepSeekopen

    57.6%
  21. 2157.0%
  22. 22
    GPT-5 Mini

    OpenAIclosed

    56.2%
  23. 23
    GPT-5 Nano

    OpenAIclosed

    54.7%
  24. 24
    Gemini 2.5 Pro

    Googleclosed

    53.6%
  25. 25
    Claude 3.7 Sonnet

    Anthropicclosed

    52.8%
  26. 26
    DeepSeek R1

    DeepSeekopen

    49.2%
  27. 2747.7%
  28. 28
    Devstral Small

    Mistralopen

    46.8%
  29. 29
    o4-mini

    OpenAIclosed

    45.0%
  30. 30
    Claude Haiku 3.5

    Anthropicclosed

    40.6%
  31. 31
    GPT-4.1

    OpenAIclosed

    39.6%
  32. 32
    Claude 3.5 Sonnet

    Anthropicclosed

    33.6%
  33. 33
    Gemini 2.5 Flash

    Googleclosed

    28.7%
  34. 34
    gpt-oss-120B

    OpenAIopen

    26.0%
  35. 35
    GPT-4.1 Mini

    OpenAIclosed

    23.9%
  36. 3621.0%
  37. 37
    Gemini 2.0 Flash

    Googleclosed

    13.5%
  38. 389.1%
  39. 399.0%

Back to all benchmarks