LLMcompare

Search

Search for a command to run...

Coding

DeepSWE

DataCurve's long-horizon software-engineering benchmark of 113 original tasks across TypeScript, Go, Python, JavaScript, and Rust with program-based verifiers; scores are agent-system results under the eval harness.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
29 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    GPT-6.1 Sol

    OpenAIclosed

    75.2%
  2. 2
    GPT-6 Astra

    OpenAIclosed

    74.1%
  3. 3
    Claude Opus 5

    Anthropicclosed

    73.7%
  4. 4
    Gemini 3.8 Flash

    Googleclosed

    73.7%
  5. 5
    GPT-5.6 Sol

    OpenAIclosed

    72.7%
  6. 6
    Claude Fable 5

    Anthropicclosed

    69.9%
  7. 7
    Kimi K3

    Moonshotopen

    69.0%
  8. 8
    GPT-6 Sol

    OpenAIclosed

    68.8%
  9. 9
    Claude Fable 5.1

    Anthropicclosed

    67.4%
  10. 10
    GPT-5.6 Luna

    OpenAIclosed

    67.0%
  11. 11
    GPT-5.5

    OpenAIclosed

    67.0%
  12. 12
    Grok 4.6

    xAIclosed

    67.0%
  13. 13
    GLM-5.3

    Zhipu AIopen

    66.9%
  14. 14
    GPT-6 Luna

    OpenAIclosed

    66.6%
  15. 15
    Gemini 3.7 Flash

    Googleclosed

    65.3%
  16. 16
    Hy4 Preview

    Tencentopen

    64.3%
  17. 17
    GLM-5.3 Flash

    Zhipu AIopen

    63.0%
  18. 18
    DeepSeek V4 Pro 0813

    DeepSeekclosed

    62.7%
  19. 19
    Claude Opus 4.8

    Anthropicclosed

    59.0%
  20. 20
    Qwen3.8 Max

    Alibabaopen

    56.6%
  21. 21
    Muse Spark 1.2

    Metaclosed

    55.0%
  22. 2254.4%
  23. 23
    Claude Sonnet 5

    Anthropicclosed

    54.0%
  24. 24
    Gemini 3.6 Flash

    Googleclosed

    47.0%
  25. 25
    GLM-5.2

    Zhipu AIopen

    44.0%
  26. 26
    Qwen3.8 27B

    Alibabaopen

    42.2%
  27. 27
    Gemini 3.5 Flash

    Googleclosed

    36.0%
  28. 28
    DeepSeek V4 Pro

    DeepSeekopen

    12.8%
  29. 29
    DeepSeek V4 Flash

    DeepSeekopen

    7.3%

Back to all benchmarks