LLMcompare

Search

Search for a command to run...

Reasoning & knowledge

AIME 2025

American Invitational Mathematics Examination 2025 — contest math problems testing multi-step reasoning; model reports may differ by answer format and sampling setup.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Reasoning & knowledge
Metric
Percent
Direction
Higher is better
Coverage
65 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Mythos 5

    Anthropicclosed

    97.0%
  2. 2
    Kimi K2.5

    Moonshotopen

    96.1%
  3. 3
    GPT-5

    OpenAIclosed

    94.6%
  4. 4
    Grok 4

    xAIclosed

    92.7%
  5. 5
    o4-mini

    OpenAIclosed

    92.7%
  6. 6
    gpt-oss-120B

    OpenAIopen

    92.5%
  7. 7
    DeepSeek V3.2

    DeepSeekopen

    92.0%
  8. 8
    gpt-oss-20B

    OpenAIopen

    91.7%
  9. 9
    GPT-5 Mini

    OpenAIclosed

    91.1%
  10. 1089.2%
  11. 1189.1%
  12. 12
    o3

    OpenAIclosed

    88.9%
  13. 13
    DeepSeek V3.1

    DeepSeekopen

    88.4%
  14. 14
    Gemini 2.5 Pro

    Googleclosed

    87.7%
  15. 15
    DeepSeek R1-0528

    DeepSeekopen

    87.5%
  16. 1686.7%
  17. 17
    GPT-5 Nano

    OpenAIclosed

    85.2%
  18. 18
    Qwen3 235B-A22B

    Alibabaopen

    81.5%
  19. 19
    Qwen3 Max

    Alibabaclosed

    80.7%
  20. 20
    o1 Preview

    OpenAIclosed

    79.3%
  21. 2173.7%
  22. 22
    Qwen3 32B

    Alibabaopen

    72.9%
  23. 2372.1%
  24. 24
    DeepSeek R1

    DeepSeekopen

    70.0%
  25. 25
    QwQ-32B

    Alibabaopen

    69.5%
  26. 2669.5%
  27. 2763.0%
  28. 28
    Phi-4 Reasoning

    Microsoftopen

    62.9%
  29. 29
    Magistral Small

    Mistralopen

    62.8%
  30. 30
    Grok 3

    xAIclosed

    58.0%
  31. 3155.7%
  32. 3253.7%
  33. 33
    Kimi K2

    Moonshotopen

    49.5%
  34. 34
    GPT-4.1

    OpenAIclosed

    46.4%
  35. 3541.3%
  36. 36
    GPT-4.1 Mini

    OpenAIclosed

    40.2%
  37. 37
    Qwen3 Coder 480B

    Alibabaopen

    39.3%
  38. 38
    Mistral Medium 3.1

    Mistralclosed

    38.3%
  39. 39
    Mistral Large 3

    Mistralopen

    38.0%
  40. 40
    Devstral 2

    Mistralopen

    36.7%
  41. 41
    Reka Flash 3

    Rekaopen

    33.7%
  42. 42
    Ministral 3 8B

    Mistralopen

    31.7%
  43. 43
    Ministral 3 14B

    Mistralopen

    30.0%
  44. 4429.0%
  45. 4527.0%
  46. 46
    DeepSeek V3

    DeepSeekopen

    26.0%
  47. 47
    GPT-4.1 Nano

    OpenAIclosed

    24.0%
  48. 48
    Ministral 3 3B

    Mistralopen

    22.0%
  49. 49
    Gemma 3 27B

    Googleopen

    20.7%
  50. 5019.3%
  51. 51
    Gemma 3 12B

    Googleopen

    18.3%
  52. 52
    Phi-4

    Microsoftopen

    18.0%
  53. 53
    GPT-4o Mini

    OpenAIclosed

    14.7%
  54. 5414.0%
  55. 55
    Command A

    Cohereclosed

    13.0%
  56. 56
    Gemma 3 4B

    Googleopen

    12.7%
  57. 5711.0%
  58. 58
    Phi-4-mini

    Microsoftopen

    6.7%
  59. 596.0%
  60. 60
    Llama 3.1 8B

    Metaopen

    4.3%
  61. 614.0%
  62. 62
    OLMo 2 32B

    Ai2open

    3.3%
  63. 633.0%
  64. 64
    Pixtral Large

    Mistralclosed

    2.3%
  65. 65
    Phi-3 Mini 3.8B

    Microsoftopen

    0.3%

Back to all benchmarks