LLMcompare

Search

Search for a command to run...

Coding

BigCodeBench

Realistic function-level coding benchmark with 1,140 tasks across 7 languages; pass@1 depends on the release and prompting.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
45 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    GPT-4o (May 2024)

    OpenAIclosed

    51.1%
  2. 2
    DeepSeek V3

    DeepSeekopen

    50.0%
  3. 349.7%
  4. 449.0%
  5. 5
    GPT-4.1 Mini

    OpenAIclosed

    48.9%
  6. 6
    GPT-4 Turbo

    OpenAIclosed

    48.2%
  7. 7
    DeepSeek Coder V2

    DeepSeekopen

    48.2%
  8. 846.9%
  9. 9
    Claude 3.5 Sonnet

    Anthropicclosed

    46.8%
  10. 10
    GPT-4o Mini

    OpenAIclosed

    46.1%
  11. 11
    Claude Haiku 3.5

    Anthropicclosed

    46.1%
  12. 1246.1%
  13. 13
    Qwen2.5 72B

    Alibabaopen

    45.8%
  14. 14
    Phi-4

    Microsoftopen

    45.5%
  15. 15
    Claude 3 Opus

    Anthropicclosed

    45.5%
  16. 1645.0%
  17. 1744.6%
  18. 1843.9%
  19. 19
    Llama 3 70B

    Metaopen

    43.6%
  20. 20
    Gemma 2 27B

    Googleopen

    42.8%
  21. 21
    Claude 3 Sonnet

    Anthropicclosed

    42.7%
  22. 22
    Codestral 22B

    Mistralopen

    41.8%
  23. 2340.7%
  24. 24
    Mixtral 8x22B

    Mistralopen

    40.6%
  25. 2539.8%
  26. 26
    Claude 3 Haiku

    Anthropicclosed

    39.4%
  27. 2738.7%
  28. 2838.1%
  29. 29
    Yi-Large

    01.AIclosed

    37.7%
  30. 30
    Phi-3 Medium 14B

    Microsoftopen

    37.6%
  31. 31
    Qwen2.5 7B

    Alibabaopen

    37.6%
  32. 3236.8%
  33. 3335.3%
  34. 34
    Gemma 2 9B

    Googleopen

    34.7%
  35. 3533.9%
  36. 3633.8%
  37. 37
    Llama 3.1 8B

    Metaopen

    32.8%
  38. 38
    Llama 3 8B

    Metaopen

    31.9%
  39. 39
    Mistral Large

    Mistralclosed

    30.0%
  40. 40
    Phi-3 Mini 3.8B

    Microsoftopen

    29.6%
  41. 4129.0%
  42. 42
    Llama 3.2 3B

    Metaopen

    23.4%
  43. 43
    Mistral 7B

    Mistralopen

    19.5%
  44. 4410.6%
  45. 45
    Llama 3.2 1B

    Metaopen

    8.2%

Back to all benchmarks