Coding
SWE-bench Multilingual
Real-world GitHub issue resolution across 300 tasks, 42 repositories, and 9 programming languages; agent scaffold and language mix matter.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Coding
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 21 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1Hy4 Preview82.9%
Tencentopen
- 2Composer 2.579.8%
Cursorclosed
- 3Qwen3.7 Max78.3%
Alibabaclosed
- 4MiniMax M2.776.5%
MiniMaxclosed
- 5Gemini 3 Flash72.7%
Googleclosed
- 6Ling-3.0 Flash72.4%
InclusionAIopen
- 7Claude Opus 4.672.0%
Anthropicclosed
- 8Seed 2.0 Pro71.7%
ByteDanceclosed
- 9Claude Opus 4.570.7%
Anthropicclosed
- 10GLM-569.7%
Zhipu AIopen
- 11Gemini 3 Pro68.7%
Googleclosed
- 12MiniMax M2.568.3%
MiniMaxopen
- 13Nemotron 3 Ultra67.7%
NVIDIAopen
- 14Kimi K2.567.3%
Moonshotopen
- 15Claude Sonnet 4.567.0%
Anthropicclosed
- 16Claude Haiku 4.564.7%
Anthropicclosed
- 17DeepSeek V3.259.0%
DeepSeekopen
- 18Kimi K247.3%
Moonshotopen
- 19Granite 4.2 30B41.9%
IBMopen
- 20GPT-5 Mini39.7%
OpenAIclosed
- 21Granite 4.2 8B30.8%
IBMopen