LLMcompare

Search

Search for a command to run...

Reasoning & knowledge

MMLU-Pro

Harder multi-task language understanding across STEM and professional domains; reported results depend on prompt, shot count, and reasoning/tool settings.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Reasoning & knowledge
Metric
Percent
Direction
Higher is better
Coverage
88 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Kimi K2.5

    Moonshotopen

    87.1%
  2. 286.8%
  3. 3
    Grok 4

    xAIclosed

    86.6%
  4. 4
    Gemini 2.5 Pro

    Googleclosed

    86.2%
  5. 5
    DeepSeek V3.2

    DeepSeekopen

    86.2%
  6. 6
    o3

    OpenAIclosed

    85.3%
  7. 7
    Gemma 4 31B

    Googleopen

    85.2%
  8. 8
    DeepSeek R1-0528

    DeepSeekopen

    85.0%
  9. 9
    DeepSeek V3.1

    DeepSeekopen

    84.8%
  10. 10
    o1 Preview

    OpenAIclosed

    84.8%
  11. 11
    Qwen3 Max

    Alibabaclosed

    84.1%
  12. 12
    o1

    OpenAIclosed

    84.1%
  13. 13
    DeepSeek R1

    DeepSeekopen

    84.0%
  14. 14
    Gemma 4 26B

    Googleopen

    82.6%
  15. 15
    Kimi K2

    Moonshotopen

    81.1%
  16. 16
    Mistral Large 3

    Mistralopen

    80.7%
  17. 17
    GPT-4.1

    OpenAIclosed

    80.6%
  18. 1880.6%
  19. 1980.5%
  20. 20
    Grok 3

    xAIclosed

    79.9%
  21. 2179.5%
  22. 22
    o3-mini

    OpenAIclosed

    79.1%
  23. 23
    Qwen3 Coder 480B

    Alibabaopen

    78.8%
  24. 2478.3%
  25. 25
    GPT-4.1 Mini

    OpenAIclosed

    78.1%
  26. 2677.6%
  27. 2777.2%
  28. 28
    QwQ-32B

    Alibabaopen

    76.4%
  29. 29
    Devstral 2

    Mistralopen

    76.2%
  30. 30
    DeepSeek V3

    DeepSeekopen

    75.9%
  31. 3175.9%
  32. 32
    Sonar Pro

    Perplexityclosed

    75.5%
  33. 33
    Claude 3.5 Sonnet

    Anthropicclosed

    75.1%
  34. 3474.3%
  35. 35
    Phi-4 Reasoning

    Microsoftopen

    74.3%
  36. 36
    o1 Mini

    OpenAIclosed

    74.2%
  37. 3774.0%
  38. 38
    GPT-4o (May 2024)

    OpenAIclosed

    74.0%
  39. 3974.0%
  40. 4073.9%
  41. 4173.3%
  42. 42
    Command A

    Cohereclosed

    71.2%
  43. 43
    Qwen2.5 72B

    Alibabaopen

    71.1%
  44. 4470.6%
  45. 45
    Phi-4

    Microsoftopen

    70.4%
  46. 46
    Pixtral Large

    Mistralclosed

    70.1%
  47. 47
    Claude 3 Opus

    Anthropicclosed

    69.6%
  48. 48
    GPT-4 Turbo

    OpenAIclosed

    69.4%
  49. 49
    Ministral 3 14B

    Mistralopen

    69.3%
  50. 5069.1%
  51. 5169.0%
  52. 5269.0%
  53. 5368.9%
  54. 54
    Sonar

    Perplexityclosed

    68.9%
  55. 55
    Mistral Medium 3.1

    Mistralclosed

    68.3%
  56. 56
    Gemma 3 27B

    Googleopen

    67.5%
  57. 57
    Reka Flash 3

    Rekaopen

    66.9%
  58. 5866.4%
  59. 59
    GPT-4.1 Nano

    OpenAIclosed

    65.7%
  60. 60
    GPT-4o Mini

    OpenAIclosed

    64.8%
  61. 61
    Ministral 3 8B

    Mistralopen

    64.2%
  62. 6263.7%
  63. 63
    Claude Haiku 3.5

    Anthropicclosed

    63.4%
  64. 64
    Gemma 3 12B

    Googleopen

    60.6%
  65. 65
    Claude 3 Sonnet

    Anthropicclosed

    57.9%
  66. 6657.2%
  67. 6756.9%
  68. 68
    Qwen2.5 7B

    Alibabaopen

    56.3%
  69. 69
    GPT-4

    OpenAIclosed

    56.2%
  70. 7054.3%
  71. 71
    Mixtral 8x22B

    Mistralopen

    53.7%
  72. 72
    Mistral Small

    Mistralclosed

    52.9%
  73. 73
    Phi-4-mini

    Microsoftopen

    52.8%
  74. 74
    Ministral 3 3B

    Mistralopen

    52.4%
  75. 75
    Mistral Large

    Mistralclosed

    51.5%
  76. 76
    OLMo 2 32B

    Ai2open

    51.1%
  77. 77
    Claude 2.1

    Anthropicclosed

    49.5%
  78. 78
    Claude 2

    Anthropicclosed

    48.6%
  79. 79
    Phi-4-multimodal

    Microsoftopen

    48.5%
  80. 80
    Llama 3.1 8B

    Metaopen

    48.3%
  81. 81
    GPT-3.5 Turbo

    OpenAIclosed

    46.2%
  82. 8244.7%
  83. 83
    Gemma 3 4B

    Googleopen

    43.6%
  84. 84
    Phi-3 Mini 3.8B

    Microsoftopen

    43.5%
  85. 85
    Gemini 1.0 Pro

    Googleclosed

    43.1%
  86. 8642.9%
  87. 87
    DBRX Instruct

    Databricksopen

    39.7%
  88. 88
    Mixtral 8x7B

    Mistralopen

    38.7%

Back to all benchmarks