LLMcompare

Search

Search for a command to run...

Coding

LiveCodeBench

Contamination-resistant coding problems collected after model training cutoffs; score depends on the LiveCodeBench release and pass@k protocol.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Percent
Direction
Higher is better
Coverage
94 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Qwen3.8 27B

    Alibabaopen

    90.3%
  2. 289.0%
  3. 3
    Claude Mythos 5

    Anthropicclosed

    88.1%
  4. 4
    DeepSeek V3.2

    DeepSeekopen

    86.2%
  5. 5
    Kimi K2.5

    Moonshotopen

    85.0%
  6. 6
    DeepSeek R1-0528

    DeepSeekopen

    84.4%
  7. 7
    Grok 4

    xAIclosed

    81.9%
  8. 8
    o3

    OpenAIclosed

    80.8%
  9. 9
    Qwen3 235B-A22B

    Alibabaopen

    80.4%
  10. 10
    Gemini 2.5 Pro

    Googleclosed

    80.1%
  11. 11
    Gemma 4 31B

    Googleopen

    80.0%
  12. 12
    Gemma 4 26B

    Googleopen

    77.1%
  13. 13
    Qwen3 Max

    Alibabaclosed

    76.7%
  14. 1475.8%
  15. 15
    Gemini 2.5 Flash

    Googleclosed

    75.1%
  16. 16
    DeepSeek V3.1

    DeepSeekopen

    74.8%
  17. 1773.2%
  18. 18
    o3-mini

    OpenAIclosed

    71.7%
  19. 1971.1%
  20. 2068.3%
  21. 21
    o1

    OpenAIclosed

    67.9%
  22. 2267.2%
  23. 23
    Qwen3 32B

    Alibabaopen

    65.7%
  24. 24
    QwQ-32B

    Alibabaopen

    63.4%
  25. 25
    Claude Opus 4

    Anthropicclosed

    62.4%
  26. 26
    Claude Sonnet 4

    Anthropicclosed

    59.4%
  27. 27
    Qwen3 Coder 480B

    Alibabaopen

    58.5%
  28. 28
    o1 Mini

    OpenAIclosed

    57.6%
  29. 2957.5%
  30. 3057.2%
  31. 3156.6%
  32. 32
    Magistral Small

    Mistralopen

    55.8%
  33. 33
    Qwen2.5 72B

    Alibabaopen

    55.5%
  34. 34
    Phi-4 Reasoning

    Microsoftopen

    53.8%
  35. 35
    Kimi K2

    Moonshotopen

    53.7%
  36. 3653.1%
  37. 3751.2%
  38. 38
    DeepSeek V3

    DeepSeekopen

    49.6%
  39. 39
    Claude 3.5 Sonnet

    Anthropicclosed

    48.7%
  40. 4048.7%
  41. 41
    GPT-4.1 Mini

    OpenAIclosed

    48.3%
  42. 42
    Mistral Large 3

    Mistralopen

    46.5%
  43. 43
    GPT-4.1

    OpenAIclosed

    45.7%
  44. 44
    Devstral 2

    Mistralopen

    44.8%
  45. 45
    Reka Flash 3

    Rekaopen

    43.5%
  46. 4643.4%
  47. 4742.6%
  48. 48
    Grok 3

    xAIclosed

    42.5%
  49. 49
    Mistral Medium 3.1

    Mistralclosed

    40.6%
  50. 5040.3%
  51. 5139.6%
  52. 52
    GPT-4 Turbo

    OpenAIclosed

    37.3%
  53. 53
    GPT-4o Mini

    OpenAIclosed

    35.5%
  54. 54
    Ministral 3 14B

    Mistralopen

    35.1%
  55. 55
    GPT-4o (May 2024)

    OpenAIclosed

    33.4%
  56. 5633.3%
  57. 5732.8%
  58. 58
    GPT-4.1 Nano

    OpenAIclosed

    32.6%
  59. 59
    Claude Haiku 3.5

    Anthropicclosed

    31.4%
  60. 60
    Ministral 3 8B

    Mistralopen

    30.3%
  61. 61
    Gemma 3 27B

    Googleopen

    29.7%
  62. 62
    Sonar

    Perplexityclosed

    29.5%
  63. 63
    Command A

    Cohereclosed

    28.7%
  64. 64
    Qwen2.5 7B

    Alibabaopen

    28.7%
  65. 65
    Claude 3 Opus

    Anthropicclosed

    27.9%
  66. 6627.7%
  67. 6727.5%
  68. 68
    Sonar Pro

    Perplexityclosed

    27.5%
  69. 69
    Pixtral Large

    Mistralclosed

    26.1%
  70. 70
    Ministral 3 3B

    Mistralopen

    24.7%
  71. 71
    Gemma 3 12B

    Googleopen

    24.6%
  72. 7223.2%
  73. 73
    Phi-4

    Microsoftopen

    23.1%
  74. 74
    Claude 3 Haiku

    Anthropicclosed

    22.5%
  75. 7521.7%
  76. 76
    Claude 2.1

    Anthropicclosed

    19.5%
  77. 7718.0%
  78. 78
    Mistral Large

    Mistralclosed

    17.8%
  79. 79
    Claude 3 Sonnet

    Anthropicclosed

    17.5%
  80. 80
    Claude 2

    Anthropicclosed

    17.1%
  81. 8116.9%
  82. 8215.8%
  83. 83
    Mixtral 8x22B

    Mistralopen

    14.8%
  84. 8414.3%
  85. 85
    Mistral Small

    Mistralclosed

    14.1%
  86. 86
    Phi-4-multimodal

    Microsoftopen

    13.1%
  87. 87
    Phi-4-mini

    Microsoftopen

    12.6%
  88. 88
    Gemma 3 4B

    Googleopen

    12.6%
  89. 89
    Llama 3.1 8B

    Metaopen

    11.6%
  90. 90
    Phi-3 Mini 3.8B

    Microsoftopen

    11.6%
  91. 91
    Gemini 1.0 Pro

    Googleclosed

    11.6%
  92. 92
    DBRX Instruct

    Databricksopen

    9.3%
  93. 93
    OLMo 2 32B

    Ai2open

    6.8%
  94. 94
    Mixtral 8x7B

    Mistralopen

    6.6%

Back to all benchmarks