LLMcompare

Search

Search for a command to run...

Coding

NL2Repo-Bench

Long-horizon repository generation: build a full installable Python library from a natural-language spec, graded by upstream pytest; harness and task release are part of the result.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
9 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    DeepSeek V4 Pro 0813

    DeepSeekclosed

    61.5%
  2. 2
    GLM-5.3

    Zhipu AIopen

    58.0%
  3. 3
    Qwen3.8 Max

    Alibabaopen

    55.9%
  4. 454.2%
  5. 5
    Qwen3.8 27B

    Alibabaopen

    42.3%
  6. 6
    MiniMax M2.7

    MiniMaxclosed

    39.8%
  7. 7
    DeepSeek V4 Flash

    DeepSeekopen

    39.4%
  8. 8
    DeepSeek V4 Pro

    DeepSeekopen

    38.5%
  9. 9
    Seed 2.0 Pro

    ByteDanceclosed

    27.9%

Back to all benchmarks