LLMcompare

Search

Search for a command to run...

Coding

SWE-bench Pro

Long-horizon repository engineering across 1,865 tasks from 41 professional codebases; leaderboard results are agent-system scores, not model-only capability scores.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
40 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Fable 5

    Anthropicclosed

    80.0%
  2. 2
    Claude Mythos 5

    Anthropicclosed

    80.0%
  3. 3
    Claude Opus 4.8

    Anthropicclosed

    69.2%
  4. 4
    Qwen3.8 Max

    Alibabaopen

    67.7%
  5. 5
    Hy4 Preview

    Tencentopen

    65.7%
  6. 6
    GPT-5.6 Sol

    OpenAIclosed

    64.6%
  7. 7
    Claude Opus 4.7

    Anthropicclosed

    64.3%
  8. 8
    GPT-5.6 Terra

    OpenAIclosed

    63.4%
  9. 9
    GPT-5.6 Luna

    OpenAIclosed

    62.7%
  10. 10
    Qwen3.8 27B

    Alibabaopen

    61.7%
  11. 11
    Atria Dawn Preview

    Shanghai AI Labopen

    59.6%
  12. 12
    MiniMax M3

    MiniMaxopen

    59.0%
  13. 13
    GPT-5.5

    OpenAIclosed

    58.6%
  14. 14
    GPT-5.4

    OpenAIclosed

    57.7%
  15. 15
    Ling-3.0 Flash

    InclusionAIopen

    56.6%
  16. 16
    MiniMax M2.7

    MiniMaxclosed

    56.2%
  17. 17
    Inkling-Small

    Thinking Machinesopen

    55.9%
  18. 18
    GPT-5.4 Mini

    OpenAIclosed

    54.4%
  19. 19
    Inkling

    Thinking Machinesopen

    54.3%
  20. 20
    Gemini 3.1 Pro

    Googleclosed

    54.2%
  21. 21
    GPT-5.4 Nano

    OpenAIclosed

    52.4%
  22. 22
    Muse Glimmer

    Metaopen

    51.2%
  23. 23
    Kimi K2.5

    Moonshotopen

    50.7%
  24. 24
    Seed 2.0 Pro

    ByteDanceclosed

    46.9%
  25. 25
    Claude Opus 4.5

    Anthropicclosed

    45.9%
  26. 26
    Claude Sonnet 4.5

    Anthropicclosed

    43.6%
  27. 27
    Claude Sonnet 4

    Anthropicclosed

    42.7%
  28. 28
    GPT-5

    OpenAIclosed

    41.8%
  29. 29
    Claude Haiku 4.5

    Anthropicclosed

    39.5%
  30. 30
    Qwen3 Coder 480B

    Alibabaopen

    38.7%
  31. 31
    Gemini 3 Flash

    Googleclosed

    34.6%
  32. 3233.3%
  33. 33
    Kimi K2

    Moonshotopen

    27.7%
  34. 34
    Qwen3 235B-A22B

    Alibabaopen

    21.4%
  35. 3519.1%
  36. 36
    gpt-oss-120B

    OpenAIopen

    16.2%
  37. 37
    DeepSeek V3.2

    DeepSeekopen

    15.6%
  38. 38
    Gemma 3 27B

    Googleopen

    11.4%
  39. 3911.2%
  40. 405.2%

Back to all benchmarks