Reasoning & knowledge
MMMU-Pro
Multimodal reasoning over 3,460 questions in six core disciplines. MMMU-Pro hardens the original MMMU in three ways: it filters out questions answerable by text-only models, expands the choice set from 4 to 10 options, and introduces a vision-only input format in which the question itself is embedded in a screenshot. Scores run well below MMMU and the two are not comparable. Artificial Analysis publishes one MMMU-Pro figure rather than separate standard and vision splits, so this value should not be read as either split on its own.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Reasoning & knowledge
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 77 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1GPT-6 Astra86.9%
OpenAIclosed
- 2Gemini 3.8 Flash85.6%
Googleclosed
- 3Gemini 3.7 Flash85.5%
Googleclosed
- 4Claude Opus 584.7%
Anthropicclosed
- 5Gemini 3.5 Flash84.3%
Googleclosed
- 6GPT-5.6 Sol83.4%
OpenAIclosed
- 7Gemini 3.6 Flash83.2%
Googleclosed
- 8Qwen3.8 Max82.8%
Alibabaopen
- 9GPT-5.6 Terra80.7%
OpenAIclosed
- 10Qwen3.7 Plus80.5%
Alibabaclosed
- 11Kimi K380.5%
Moonshotopen
- 12Grok 4.580.4%
xAIclosed
- 13GPT-5.579.9%
OpenAIclosed
- 14Gemini 3 Flash79.9%
Googleclosed
- 15Qwen3.8 Flash Next79.8%
Alibabaopen
- 16Kimi K2.679.4%
Moonshotclosed
- 17Gemini 3.5 Flash-Lite79.0%
Googleclosed
- 18Claude Opus 4.778.8%
Anthropicclosed
- 19GPT-5.6 Luna78.6%
OpenAIclosed
- 20MiniMax M378.6%
MiniMaxopen
- 21GPT-5.478.4%
OpenAIclosed
- 22Grok 4.378.1%
xAIclosed
- 23Qwen3.6 Plus78.0%
Alibabaclosed
- 24Claude Sonnet 577.3%
Anthropicclosed
- 25Qwen3.5 397B-A17B77.3%
Alibabaopen
- 26DeepSeek V4.1 Flash77.0%
DeepSeekopen
- 27Qwen3.8 27B76.3%
Alibabaopen
- 28Gemini 3.1 Flash-Lite75.5%
Googleclosed
- 29MiMo-V2.575.4%
Xiaomiopen
- 30Kimi K2.575.4%
Moonshotopen
- 31Claude Opus 4.675.4%
Anthropicclosed
- 32Qwen3.6 35B A3B75.0%
Alibabaopen
- 33Gemini 2.5 Pro74.9%
Googleclosed
- 34Qwen3.6 27B74.6%
Alibabaopen
- 35Muse Glimmer74.3%
Metaopen
- 36GPT-574.2%
OpenAIclosed
- 37Claude Opus 4.574.0%
Anthropicclosed
- 38Inkling-Small74.0%
Thinking Machinesopen
- 39Inkling73.5%
Thinking Machinesopen
- 40Gemma 4 31B73.4%
Googleopen
- 41Claude Sonnet 4.673.3%
Anthropicclosed
- 42GPT-5.4 Mini73.3%
OpenAIclosed
- 43GPT-5 Mini70.1%
OpenAIclosed
- 44o370.1%
OpenAIclosed
- 45o4-mini69.2%
OpenAIclosed
- 46Gemini 2.5 Flash69.1%
Googleclosed
- 47Grok 468.8%
xAIclosed
- 48GPT-5.4 Nano65.4%
OpenAIclosed
- 49Mistral Medium 3.564.9%
Mistralopen
- 50Llama 4 Maverick62.1%
Metaopen
- 51Grok 4 Fast61.8%
xAIclosed
- 52GPT-4.161.2%
OpenAIclosed
- 53GPT-5 Nano61.0%
OpenAIclosed
- 54GPT-4.1 Mini58.7%
OpenAIclosed
- 55Gemini 2.5 Flash-Lite58.2%
Googleclosed
- 56Mistral Small 456.8%
Mistralopen
- 57Mistral Large 355.7%
Mistralopen
- 58Qwen3-Omni-30B-A3B-Instruct55.5%
Alibabaopen
- 59Gemini 1.5 Pro55.0%
Googleclosed
- 60Mistral Medium 3.154.2%
Mistralclosed
- 61Llama 4 Scout52.9%
Metaopen
- 62Pixtral Large50.6%
Mistralclosed
- 63Ministral 3 14B49.8%
Mistralopen
- 64Gemini 1.5 Flash48.4%
Googleclosed
- 65Mistral Small 3.248.0%
Mistralopen
- 66Gemma 3 27B48.0%
Googleopen
- 67Ministral 3 8B46.0%
Mistralopen
- 68GPT-4o Mini41.5%
OpenAIclosed
- 69GPT-4.1 Nano40.1%
OpenAIclosed
- 70Llama 3.2 90B Vision39.5%
Metaopen
- 71Ministral 3 3B38.1%
Mistralopen
- 72Gemma 3 12B37.5%
Googleopen
- 73Gemini 1.5 Flash-8B36.5%
Googleclosed
- 74Claude 3 Haiku30.8%
Anthropicclosed
- 75Gemma 3 4B29.9%
Googleopen
- 76Llama 3.2 11B Vision29.3%
Metaopen
- 77Phi-4-multimodal14.5%
Microsoftopen