LLMcompare

Search

Search for a command to run...

Composite indices

Artificial Analysis Intelligence Index v4.3.2

Artificial Analysis composite across AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. Keep separate from prior index versions because its evaluation components changed.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Composite indices
Metric
Index (0-100)
Direction
Higher is better
Coverage
8 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Opus 5.5

    Anthropicclosed

    58.0
  2. 2
    Claude Sonnet 5.5

    Anthropicclosed

    56.0
  3. 3
    GPT-6 Astra

    OpenAIclosed

    53.0
  4. 4
    Claude Fable 5.1

    Anthropicclosed

    53.0
  5. 5
    GPT-6.1 Sol

    OpenAIclosed

    52.0
  6. 6
    GPT-6 Sol

    OpenAIclosed

    48.0
  7. 7
    Grok 4.7

    xAIclosed

    46.0
  8. 8
    GPT-6 Luna

    OpenAIclosed

    37.0

Back to all benchmarks