LLMcompare

Search

Search for a command to run...

Coding

SWE-bench Multilingual

Real-world GitHub issue resolution across 300 tasks, 42 repositories, and 9 programming languages; agent scaffold and language mix matter.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
21 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Hy4 Preview

    Tencentopen

    82.9%
  2. 2
    Composer 2.5

    Cursorclosed

    79.8%
  3. 3
    Qwen3.7 Max

    Alibabaclosed

    78.3%
  4. 4
    MiniMax M2.7

    MiniMaxclosed

    76.5%
  5. 5
    Gemini 3 Flash

    Googleclosed

    72.7%
  6. 6
    Ling-3.0 Flash

    InclusionAIopen

    72.4%
  7. 7
    Claude Opus 4.6

    Anthropicclosed

    72.0%
  8. 8
    Seed 2.0 Pro

    ByteDanceclosed

    71.7%
  9. 9
    Claude Opus 4.5

    Anthropicclosed

    70.7%
  10. 10
    GLM-5

    Zhipu AIopen

    69.7%
  11. 11
    Gemini 3 Pro

    Googleclosed

    68.7%
  12. 12
    MiniMax M2.5

    MiniMaxopen

    68.3%
  13. 1367.7%
  14. 14
    Kimi K2.5

    Moonshotopen

    67.3%
  15. 15
    Claude Sonnet 4.5

    Anthropicclosed

    67.0%
  16. 16
    Claude Haiku 4.5

    Anthropicclosed

    64.7%
  17. 17
    DeepSeek V3.2

    DeepSeekopen

    59.0%
  18. 18
    Kimi K2

    Moonshotopen

    47.3%
  19. 1941.9%
  20. 20
    GPT-5 Mini

    OpenAIclosed

    39.7%
  21. 2130.8%

Back to all benchmarks