LLMcompare

Search

Search for a command to run...

Composite indices

Artificial Analysis Intelligence Index v4.3

Artificial Analysis's own composite metric on a 0-100 scale, not a percentage. v4.3 is a weighted average of ten independently-run evaluations across four equally-weighted (25% each) categories: agents (AA-Briefcase, GDPval-AA v2, AutomationBench-AA), coding (Terminal-Bench 4.0, SciCode), general (AA-Omniscience, GDP.pdf, AA-LCR v1.1), and scientific reasoning (Humanity's Last Exam, CritPt). Relative to v4.2, v4.3 removes τ³-Banking, adds AutomationBench-AA, and swaps in Terminal-Bench 4.0 for Terminal-Bench 2.1. Index versions are not comparable to each other, so v4.3 scores must not be merged with v4.2 scores; several components are agent-harness evaluations, so the index is not a model-only capability measurement. Only listings Artificial Analysis has fully measured are recorded here; its estimated index values are omitted.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Composite indices
Metric
Index (0-100)
Direction
Higher is better
Coverage
88 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Fable 5.1

    Anthropicclosed

    53.4
  2. 2
    GPT-6 Astra

    OpenAIclosed

    52.8
  3. 3
    Claude Opus 5

    Anthropicclosed

    50.7
  4. 4
    Claude Fable 5

    Anthropicclosed

    49.7
  5. 5
    Muse Spark 1.3

    Metaclosed

    48.2
  6. 6
    GPT-5.6 Sol

    OpenAIclosed

    47.1
  7. 7
    Qwen3.8 Max

    Alibabaopen

    45.4
  8. 8
    GLM-5.3

    Zhipu AIopen

    44.9
  9. 9
    Grok 4.6

    xAIclosed

    44.4
  10. 10
    Kimi K3

    Moonshotopen

    43.8
  11. 11
    GPT-5.6 Terra

    OpenAIclosed

    42.3
  12. 12
    Claude Opus 4.8

    Anthropicclosed

    42.0
  13. 13
    GLM-5.3 Flash

    Zhipu AIopen

    41.9
  14. 14
    Gemini 3.8 Flash

    Googleclosed

    41.2
  15. 1539.9
  16. 16
    Muse Spark 1.2

    Metaclosed

    39.8
  17. 1739.5
  18. 18
    Gemini 3.7 Flash

    Googleclosed

    39.4
  19. 19
    Grok 4.5

    xAIclosed

    39.1
  20. 20
    GPT-5.5

    OpenAIclosed

    38.6
  21. 21
    Claude Sonnet 5

    Anthropicclosed

    38.4
  22. 22
    GPT-5.6 Luna

    OpenAIclosed

    37.5
  23. 23
    DeepSeek V4 Pro 0813

    DeepSeekclosed

    36.3
  24. 2434.5
  25. 25
    Gemini 3.6 Flash

    Googleclosed

    34.3
  26. 26
    Muse Spark 1.1

    Metaclosed

    34.3
  27. 27
    GLM-5.2

    Zhipu AIopen

    34.0
  28. 28
    Qwen3.8 27B

    Alibabaopen

    33.9
  29. 29
    Gemini 3.5 Flash

    Googleclosed

    33.0
  30. 30
    DeepSeek V4 Pro

    DeepSeekopen

    30.9
  31. 31
    K2 Horizon 375B A23B

    MBZUAI Institute of Foundation Modelsopen

    30.8
  32. 32
    Claude Sonnet 4.6

    Anthropicclosed

    30.5
  33. 33
    Gemini 3.1 Pro

    Googleclosed

    30.4
  34. 34
    Qwen3.7 Max

    Alibabaclosed

    29.9
  35. 35
    MiniMax M3

    MiniMaxopen

    29.6
  36. 36
    Kimi K2.6

    Moonshotclosed

    27.5
  37. 37
    MiMo-V2.5-Pro

    Xiaomiopen

    26.4
  38. 38
    GLM-5.1

    Zhipu AIopen

    26.4
  39. 39
    Kimi K2.7 Code

    Moonshotopen

    26.3
  40. 40
    Inkling-Small

    Thinking Machinesopen

    26.1
  41. 41
    Qwen3.7 Plus

    Alibabaclosed

    25.8
  42. 42
    Hy3

    Tencentopen

    25.8
  43. 43
    Inkling

    Thinking Machinesopen

    25.5
  44. 44
    Grok 4.3

    xAIclosed

    25.4
  45. 45
    DeepSeek V4 Flash

    DeepSeekopen

    24.6
  46. 46
    GPT-5.4 Mini

    OpenAIclosed

    24.6
  47. 4723.4
  48. 48
    MiniMax M2.7

    MiniMaxclosed

    23.2
  49. 4922.7
  50. 50
    MiMo-V2.5

    Xiaomiopen

    22.3
  51. 51
    Qwen3.6 27B

    Alibabaopen

    21.9
  52. 52
    GPT-5.4 Nano

    OpenAIclosed

    21.2
  53. 53
    Ling-3.0 Flash

    InclusionAIopen

    20.6
  54. 54
    LongCat-2.0

    Meituanopen

    19.7
  55. 5519.1
  56. 56
    Qwen3.6 35B A3B

    Alibabaopen

    18.8
  57. 57
    Muse Glimmer

    Metaopen

    18.1
  58. 58
    Claude Haiku 4.5

    Anthropicclosed

    17.6
  59. 59
    GPT-5 Mini

    OpenAIclosed

    17.4
  60. 60
    Ring-2.6-1T

    InclusionAIopen

    17.3
  61. 61
    Gemini 2.5 Pro

    Googleclosed

    16.7
  62. 6216.0
  63. 63
    Gemma 4 31B

    Googleopen

    15.4
  64. 6414.9
  65. 6513.6
  66. 66
    gpt-oss-120B

    OpenAIopen

    12.3
  67. 67
    Ling-3.0 Tiny

    InclusionAIopen

    11.9
  68. 6811.8
  69. 69
    Mistral Small 4

    Mistralopen

    11.5
  70. 70
    Mercury 2

    Inception Labsclosed

    11.5
  71. 71
    DeepSeek R1

    DeepSeekopen

    11.4
  72. 7210.9
  73. 73
    Mistral Large 3

    Mistralopen

    9.7
  74. 74
    Mistral Medium 3.1

    Mistralclosed

    9.5
  75. 75
    Devstral 2

    Mistralopen

    9.4
  76. 769.3
  77. 779.1
  78. 78
    gpt-oss-20B

    OpenAIopen

    9.0
  79. 79
    DeepSeek V3

    DeepSeekopen

    8.5
  80. 80
    Solar Pro 3

    Upstageopen

    7.8
  81. 81
    Qwen3 32B

    Alibabaopen

    7.2
  82. 827.0
  83. 836.5
  84. 84
    Ministral 3 14B

    Mistralopen

    6.0
  85. 85
    Ministral 3 8B

    Mistralopen

    5.5
  86. 86
    Gemma 3 27B

    Googleopen

    4.9
  87. 87
    Ministral 3 3B

    Mistralopen

    4.8
  88. 88
    Gemma 3 12B

    Googleopen

    3.8

Back to all benchmarks