LLMcompare

Search

Search for a command to run...

Composite indices

Artificial Analysis Intelligence Index v4.2

Artificial Analysis's own composite metric on a 0-100 scale, not a percentage. v4.2 is a weighted average of ten independently-run evaluations across four categories: agents 30% (AA-Briefcase 15%, GDPval-AA v2 10%, τ³-Banking 5%), coding 20% (Terminal-Bench 2.1 10%, SciCode 10%), general 30% (AA-Omniscience 15%, split 10% accuracy and 5% non-hallucination, GDP.pdf 10%, AA-LCR v1.1 5%), and scientific reasoning 20% (Humanity's Last Exam 10%, CritPt 10%). Index versions are not comparable to each other — v4.2 dropped GPQA Diamond and added AA-Briefcase and GDP.pdf relative to v4.1.1 — and several components are agent-harness evaluations, so the index is not a model-only capability measurement. Only listings Artificial Analysis has fully measured are recorded here; its estimated index values are omitted.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Composite indices
Metric
Index (0-100)
Direction
Higher is better
Coverage
87 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Fable 5.1

    Anthropicclosed

    56.8
  2. 2
    GPT-6 Astra

    OpenAIclosed

    54.7
  3. 3
    Claude Opus 5

    Anthropicclosed

    54.1
  4. 4
    Claude Fable 5

    Anthropicclosed

    53.2
  5. 5
    Muse Spark 1.3

    Metaclosed

    53.0
  6. 6
    GPT-5.6 Sol

    OpenAIclosed

    51.3
  7. 7
    Grok 4.6

    xAIclosed

    50.6
  8. 8
    Kimi K3

    Moonshotopen

    50.2
  9. 9
    GLM-5.3

    Zhipu AIopen

    48.6
  10. 10
    Gemini 3.8 Flash

    Googleclosed

    47.1
  11. 11
    Qwen3.8 Max

    Alibabaopen

    46.9
  12. 12
    GPT-5.6 Terra

    OpenAIclosed

    46.8
  13. 13
    Muse Spark 1.2

    Metaclosed

    46.8
  14. 14
    GLM-5.3 Flash

    Zhipu AIopen

    46.2
  15. 15
    Grok 4.5

    xAIclosed

    45.5
  16. 16
    Gemini 3.7 Flash

    Googleclosed

    45.2
  17. 17
    Claude Sonnet 5

    Anthropicclosed

    45.1
  18. 18
    GPT-5.6 Luna

    OpenAIclosed

    43.4
  19. 19
    DeepSeek V4 Pro 0813

    DeepSeekclosed

    42.1
  20. 20
    Claude Opus 4.8

    Anthropicclosed

    42.0
  21. 21
    Qwen3.8 27B

    Alibabaopen

    41.4
  22. 2240.8
  23. 23
    Gemini 3.6 Flash

    Googleclosed

    40.3
  24. 2439.9
  25. 25
    Gemini 3.5 Flash

    Googleclosed

    39.7
  26. 2639.5
  27. 27
    GPT-5.5

    OpenAIclosed

    38.6
  28. 28
    K2 Horizon 375B A23B

    MBZUAI Institute of Foundation Modelsopen

    37.8
  29. 29
    Gemini 3.1 Pro

    Googleclosed

    36.7
  30. 30
    MiniMax M3

    MiniMaxopen

    35.7
  31. 31
    Muse Spark 1.1

    Metaclosed

    34.3
  32. 32
    GLM-5.2

    Zhipu AIopen

    34.0
  33. 33
    MiMo-V2.5-Pro

    Xiaomiopen

    32.6
  34. 34
    Inkling

    Thinking Machinesopen

    32.2
  35. 35
    DeepSeek V4 Pro

    DeepSeekopen

    30.9
  36. 36
    Claude Sonnet 4.6

    Anthropicclosed

    30.5
  37. 37
    Qwen3.7 Max

    Alibabaclosed

    29.9
  38. 3827.6
  39. 39
    Kimi K2.6

    Moonshotclosed

    27.5
  40. 40
    Ling-3.0 Flash

    InclusionAIopen

    27.4
  41. 41
    GLM-5.1

    Zhipu AIopen

    26.4
  42. 42
    Kimi K2.7 Code

    Moonshotopen

    26.3
  43. 43
    Inkling-Small

    Thinking Machinesopen

    26.1
  44. 44
    Qwen3.7 Plus

    Alibabaclosed

    25.8
  45. 45
    Hy3

    Tencentopen

    25.8
  46. 46
    Grok 4.3

    xAIclosed

    25.4
  47. 47
    DeepSeek V4 Flash

    DeepSeekopen

    24.6
  48. 48
    GPT-5.4 Mini

    OpenAIclosed

    24.6
  49. 49
    Muse Glimmer

    Metaopen

    24.4
  50. 5023.4
  51. 51
    MiniMax M2.7

    MiniMaxclosed

    23.2
  52. 52
    Claude Haiku 4.5

    Anthropicclosed

    22.5
  53. 53
    MiMo-V2.5

    Xiaomiopen

    22.3
  54. 54
    Qwen3.6 27B

    Alibabaopen

    21.9
  55. 55
    GPT-5.4 Nano

    OpenAIclosed

    21.2
  56. 56
    LongCat-2.0

    Meituanopen

    19.7
  57. 5719.1
  58. 58
    Qwen3.6 35B A3B

    Alibabaopen

    18.8
  59. 59
    GPT-5 Mini

    OpenAIclosed

    17.4
  60. 60
    Ring-2.6-1T

    InclusionAIopen

    17.3
  61. 61
    Gemini 2.5 Pro

    Googleclosed

    16.7
  62. 6216.0
  63. 63
    Ling-3.0 Tiny

    InclusionAIopen

    15.7
  64. 64
    gpt-oss-120B

    OpenAIopen

    15.6
  65. 65
    Gemma 4 31B

    Googleopen

    15.4
  66. 6614.9
  67. 6713.9
  68. 6813.6
  69. 69
    Mistral Small 4

    Mistralopen

    11.5
  70. 70
    Mercury 2

    Inception Labsclosed

    11.5
  71. 71
    DeepSeek R1

    DeepSeekopen

    11.4
  72. 72
    Mistral Large 3

    Mistralopen

    11.1
  73. 7310.9
  74. 7410.7
  75. 75
    Devstral 2

    Mistralopen

    9.4
  76. 769.3
  77. 77
    gpt-oss-20B

    OpenAIopen

    9.0
  78. 78
    DeepSeek V3

    DeepSeekopen

    8.5
  79. 79
    Solar Pro 3

    Upstageopen

    7.8
  80. 80
    Qwen3 32B

    Alibabaopen

    7.2
  81. 817.0
  82. 826.5
  83. 83
    Ministral 3 14B

    Mistralopen

    6.0
  84. 84
    Ministral 3 8B

    Mistralopen

    5.5
  85. 85
    Gemma 3 27B

    Googleopen

    4.9
  86. 86
    Ministral 3 3B

    Mistralopen

    4.8
  87. 87
    Gemma 3 12B

    Googleopen

    3.8

Back to all benchmarks