Coding
NL2Repo-Bench
Long-horizon repository generation: build a full installable Python library from a natural-language spec, graded by upstream pytest; harness and task release are part of the result.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Coding
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 9 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1DeepSeek V4 Pro 081361.5%
DeepSeekclosed
- 2GLM-5.358.0%
Zhipu AIopen
- 3Qwen3.8 Max55.9%
Alibabaopen
- 4DeepSeek V4 Flash 073154.2%
DeepSeekopen
- 5Qwen3.8 27B42.3%
Alibabaopen
- 6MiniMax M2.739.8%
MiniMaxclosed
- 7DeepSeek V4 Flash39.4%
DeepSeekopen
- 8DeepSeek V4 Pro38.5%
DeepSeekopen
- 9Seed 2.0 Pro27.9%
ByteDanceclosed