LLMcompare

Search

Search for a command to run...

Reasoning & knowledge

MATH-500

500-problem subset of the MATH competition dataset covering algebra, geometry, and number theory; scores are protocol-sensitive and not interchangeable with AIME results.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Reasoning & knowledge
Metric
Percent
Direction
Higher is better
Coverage
67 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    o3

    OpenAIclosed

    99.2%
  2. 2
    Grok 4

    xAIclosed

    99.0%
  3. 3
    DeepSeek R1-0528

    DeepSeekopen

    98.3%
  4. 4
    Qwen3 235B-A22B

    Alibabaopen

    98.0%
  5. 5
    QwQ-32B

    Alibabaopen

    98.0%
  6. 697.8%
  7. 7
    Kimi K2

    Moonshotopen

    97.4%
  8. 8
    o3-mini

    OpenAIclosed

    97.3%
  9. 9
    DeepSeek R1

    DeepSeekopen

    97.3%
  10. 10
    Qwen3 32B

    Alibabaopen

    97.2%
  11. 11
    o1

    OpenAIclosed

    97.0%
  12. 12
    Gemini 2.5 Pro

    Googleclosed

    96.7%
  13. 13
    Sonar Reasoning Pro

    Perplexityclosed

    95.7%
  14. 1494.5%
  15. 15
    o1 Mini

    OpenAIclosed

    94.4%
  16. 1694.3%
  17. 17
    Qwen3 Coder 480B

    Alibabaopen

    94.2%
  18. 1893.9%
  19. 19
    GPT-4.1 Mini

    OpenAIclosed

    92.5%
  20. 20
    o1 Preview

    OpenAIclosed

    92.4%
  21. 21
    Sonar Reasoning

    Perplexityclosed

    92.1%
  22. 22
    GPT-4.1

    OpenAIclosed

    91.3%
  23. 23
    DeepSeek V3

    DeepSeekopen

    90.2%
  24. 2489.3%
  25. 25
    Reka Flash 3

    Rekaopen

    89.3%
  26. 2689.1%
  27. 2788.9%
  28. 2888.3%
  29. 29
    Gemma 3 27B

    Googleopen

    88.3%
  30. 30
    Grok 3

    xAIclosed

    87.0%
  31. 31
    Gemma 3 12B

    Googleopen

    85.3%
  32. 32
    GPT-4.1 Nano

    OpenAIclosed

    84.8%
  33. 3384.4%
  34. 34
    Command A

    Cohereclosed

    81.9%
  35. 35
    Sonar

    Perplexityclosed

    81.7%
  36. 36
    Phi-4

    Microsoftopen

    81.0%
  37. 37
    GPT-4o (May 2024)

    OpenAIclosed

    79.1%
  38. 38
    GPT-4o Mini

    OpenAIclosed

    78.9%
  39. 3977.1%
  40. 40
    Gemma 3 4B

    Googleopen

    76.6%
  41. 41
    Sonar Pro

    Perplexityclosed

    74.5%
  42. 42
    DeepSeek Coder V2

    DeepSeekopen

    74.3%
  43. 43
    GPT-4 Turbo

    OpenAIclosed

    73.7%
  44. 4473.3%
  45. 45
    Claude Haiku 3.5

    Anthropicclosed

    72.1%
  46. 46
    Pixtral Large

    Mistralclosed

    71.4%
  47. 4770.3%
  48. 48
    Phi-4-mini

    Microsoftopen

    69.6%
  49. 49
    Claude 3.5 Sonnet

    Anthropicclosed

    69.5%
  50. 50
    Phi-4-multimodal

    Microsoftopen

    69.3%
  51. 5168.9%
  52. 5264.9%
  53. 53
    Claude 3 Opus

    Anthropicclosed

    64.1%
  54. 5460.6%
  55. 55
    GPT-4

    OpenAIclosed

    56.8%
  56. 56
    Mistral Small

    Mistralclosed

    56.3%
  57. 57
    Mixtral 8x22B

    Mistralopen

    54.5%
  58. 58
    Mistral Large

    Mistralclosed

    52.7%
  59. 59
    Llama 3.1 8B

    Metaopen

    51.9%
  60. 60
    Phi-3 Mini 3.8B

    Microsoftopen

    45.7%
  61. 61
    GPT-3.5 Turbo

    OpenAIclosed

    44.1%
  62. 62
    Claude 3 Sonnet

    Anthropicclosed

    41.4%
  63. 63
    Gemini 1.0 Pro

    Googleclosed

    40.3%
  64. 64
    Claude 3 Haiku

    Anthropicclosed

    39.4%
  65. 65
    Claude 2.1

    Anthropicclosed

    37.4%
  66. 66
    Mixtral 8x7B

    Mistralopen

    29.9%
  67. 67
    DBRX Instruct

    Databricksopen

    27.9%

Back to all benchmarks