LLMcompare

Search

Search for a command to run...

Reasoning & knowledge

MMMU-Pro

Multimodal reasoning over 3,460 questions in six core disciplines. MMMU-Pro hardens the original MMMU in three ways: it filters out questions answerable by text-only models, expands the choice set from 4 to 10 options, and introduces a vision-only input format in which the question itself is embedded in a screenshot. Scores run well below MMMU and the two are not comparable. Artificial Analysis publishes one MMMU-Pro figure rather than separate standard and vision splits, so this value should not be read as either split on its own.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Reasoning & knowledge
Metric
Percent
Direction
Higher is better
Coverage
77 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    GPT-6 Astra

    OpenAIclosed

    86.9%
  2. 2
    Gemini 3.8 Flash

    Googleclosed

    85.6%
  3. 3
    Gemini 3.7 Flash

    Googleclosed

    85.5%
  4. 4
    Claude Opus 5

    Anthropicclosed

    84.7%
  5. 5
    Gemini 3.5 Flash

    Googleclosed

    84.3%
  6. 6
    GPT-5.6 Sol

    OpenAIclosed

    83.4%
  7. 7
    Gemini 3.6 Flash

    Googleclosed

    83.2%
  8. 8
    Qwen3.8 Max

    Alibabaopen

    82.8%
  9. 9
    GPT-5.6 Terra

    OpenAIclosed

    80.7%
  10. 10
    Qwen3.7 Plus

    Alibabaclosed

    80.5%
  11. 11
    Kimi K3

    Moonshotopen

    80.5%
  12. 12
    Grok 4.5

    xAIclosed

    80.4%
  13. 13
    GPT-5.5

    OpenAIclosed

    79.9%
  14. 14
    Gemini 3 Flash

    Googleclosed

    79.9%
  15. 1579.8%
  16. 16
    Kimi K2.6

    Moonshotclosed

    79.4%
  17. 1779.0%
  18. 18
    Claude Opus 4.7

    Anthropicclosed

    78.8%
  19. 19
    GPT-5.6 Luna

    OpenAIclosed

    78.6%
  20. 20
    MiniMax M3

    MiniMaxopen

    78.6%
  21. 21
    GPT-5.4

    OpenAIclosed

    78.4%
  22. 22
    Grok 4.3

    xAIclosed

    78.1%
  23. 23
    Qwen3.6 Plus

    Alibabaclosed

    78.0%
  24. 24
    Claude Sonnet 5

    Anthropicclosed

    77.3%
  25. 2577.3%
  26. 2677.0%
  27. 27
    Qwen3.8 27B

    Alibabaopen

    76.3%
  28. 2875.5%
  29. 29
    MiMo-V2.5

    Xiaomiopen

    75.4%
  30. 30
    Kimi K2.5

    Moonshotopen

    75.4%
  31. 31
    Claude Opus 4.6

    Anthropicclosed

    75.4%
  32. 32
    Qwen3.6 35B A3B

    Alibabaopen

    75.0%
  33. 33
    Gemini 2.5 Pro

    Googleclosed

    74.9%
  34. 34
    Qwen3.6 27B

    Alibabaopen

    74.6%
  35. 35
    Muse Glimmer

    Metaopen

    74.3%
  36. 36
    GPT-5

    OpenAIclosed

    74.2%
  37. 37
    Claude Opus 4.5

    Anthropicclosed

    74.0%
  38. 38
    Inkling-Small

    Thinking Machinesopen

    74.0%
  39. 39
    Inkling

    Thinking Machinesopen

    73.5%
  40. 40
    Gemma 4 31B

    Googleopen

    73.4%
  41. 41
    Claude Sonnet 4.6

    Anthropicclosed

    73.3%
  42. 42
    GPT-5.4 Mini

    OpenAIclosed

    73.3%
  43. 43
    GPT-5 Mini

    OpenAIclosed

    70.1%
  44. 44
    o3

    OpenAIclosed

    70.1%
  45. 45
    o4-mini

    OpenAIclosed

    69.2%
  46. 46
    Gemini 2.5 Flash

    Googleclosed

    69.1%
  47. 47
    Grok 4

    xAIclosed

    68.8%
  48. 48
    GPT-5.4 Nano

    OpenAIclosed

    65.4%
  49. 4964.9%
  50. 5062.1%
  51. 51
    Grok 4 Fast

    xAIclosed

    61.8%
  52. 52
    GPT-4.1

    OpenAIclosed

    61.2%
  53. 53
    GPT-5 Nano

    OpenAIclosed

    61.0%
  54. 54
    GPT-4.1 Mini

    OpenAIclosed

    58.7%
  55. 5558.2%
  56. 56
    Mistral Small 4

    Mistralopen

    56.8%
  57. 57
    Mistral Large 3

    Mistralopen

    55.7%
  58. 5855.5%
  59. 59
    Gemini 1.5 Pro

    Googleclosed

    55.0%
  60. 60
    Mistral Medium 3.1

    Mistralclosed

    54.2%
  61. 6152.9%
  62. 62
    Pixtral Large

    Mistralclosed

    50.6%
  63. 63
    Ministral 3 14B

    Mistralopen

    49.8%
  64. 64
    Gemini 1.5 Flash

    Googleclosed

    48.4%
  65. 6548.0%
  66. 66
    Gemma 3 27B

    Googleopen

    48.0%
  67. 67
    Ministral 3 8B

    Mistralopen

    46.0%
  68. 68
    GPT-4o Mini

    OpenAIclosed

    41.5%
  69. 69
    GPT-4.1 Nano

    OpenAIclosed

    40.1%
  70. 7039.5%
  71. 71
    Ministral 3 3B

    Mistralopen

    38.1%
  72. 72
    Gemma 3 12B

    Googleopen

    37.5%
  73. 7336.5%
  74. 74
    Claude 3 Haiku

    Anthropicclosed

    30.8%
  75. 75
    Gemma 3 4B

    Googleopen

    29.9%
  76. 7629.3%
  77. 77
    Phi-4-multimodal

    Microsoftopen

    14.5%

Back to all benchmarks