LLMcompare

Search

Search for a command to run...

Benchmarks

What each evaluation measures, and the top models in this catalog by that score. Rankings use reported numbers only — models without a score for a benchmark are omitted. Data as of .

Full catalogCompare two models

LLMcompare only

Full ranking →

Full information

An LLMcompare-computed index, not a published benchmark. Eligible benches are z-scored against models released in the past year. The index is a shrunken mean of those z-scores — sum(z) / (published count + k), with k dummy average results — so a new model that is far above peers on every measured bench is not treated as average on the rest, while a short brochure set still cannot dominate. Beating a long tail of older models does not inflate the index. Saturated and thinly scored benches are left out. Models with fewer than three published eligible scores are unranked, and every contributing score stays visible with its own provenance.

224 models scored · top 10 shown

  1. 1
    Claude Fable 5.1

    Anthropicclosed

    +1.01 · 16/25
  2. 2
    GPT-6 Astra

    OpenAIclosed

    +0.88 · 13/25
  3. 3
    Claude Opus 5

    Anthropicclosed

    +0.87 · 19/25
  4. 4
    Claude Fable 5

    Anthropicclosed

    +0.74 · 20/25
  5. 5
    Claude Opus 5.5

    Anthropicclosed

    +0.70 · 5/25
  6. 6
    Muse Spark 1.3

    Metaclosed

    +0.69 · 9/25
  7. 7
    GPT-5.6 Sol

    OpenAIclosed

    +0.68 · 21/25
  8. 8
    Kimi K3

    Moonshotopen

    +0.60 · 15/25
  9. 9
    Claude Mythos 5

    Anthropicclosed

    +0.57 · 8/25
  10. 10
    Muse Spark 1.1

    Metaclosed

    +0.54 · 11/25

Coverage differs between models: a model needs 3 published eligible scores to appear at all. The index z-scores against models from the past year, then takes a shrunken mean of those z-scores so a new model that leads every measured bench is not treated as average on the rest. The count next to each score shows how many of the 25 eligible benches the model actually has. The published benchmarks below are the source of truth.

Reasoning & knowledge

Composite ranking →

Graduate-level Google-proof Q&A in biology, physics, and chemistry (diamond subset); scores are only comparable when evaluation protocol and tool access match.

215 models scored · top 10 shown

Full informationsource

Expert-level closed-ended questions across dozens of academic fields; tool access, modality, and evaluation protocol must match before scores are compared.

212 models scored · top 10 shown

American Invitational Mathematics Examination 2025 — contest math problems testing multi-step reasoning; model reports may differ by answer format and sampling setup.

65 models scored · top 10 shown

500-problem subset of the MATH competition dataset covering algebra, geometry, and number theory; scores are protocol-sensitive and not interchangeable with AIME results.

67 models scored · top 10 shown

Complex Research using Integrated Thinking — Physics Test: 71 unpublished research-level physics challenges (decomposed into 190 checkpoint tasks) across 11 subfields, authored by 50+ working physicists. Reported here as Artificial Analysis's independently-run accuracy on the composite challenges; tool access materially changes results, so runs with and without code execution are not comparable.

99 models scored · top 10 shown

Artificial Analysis Long Context Reasoning: 100 open-answer questions requiring multi-step synthesis across documents of 10k-100k tokens (academic papers, financial reports, legal and government documents). Pass/fail graded by an LLM judge against official answers; scores are tied to the AA-LCR version and grader model and are not a needle-in-a-haystack retrieval measure.

151 models scored · top 10 shown

Full informationsource

Share of AA-Omniscience's 6,000 factual-recall questions answered correctly, across 42 topics in six domains (business, humanities and social sciences, STEM, health, law, and software engineering). This is the accuracy metric only — keep it distinct from the AA-Omniscience Index, a separate -100 to 100 metric that penalises hallucination and rewards abstention, and from the hallucination rate.

163 models scored · top 10 shown

Multimodal reasoning over 3,460 questions in six core disciplines. MMMU-Pro hardens the original MMMU in three ways: it filters out questions answerable by text-only models, expands the choice set from 4 to 10 options, and introduces a vision-only input format in which the question itself is embedded in a screenshot. Scores run well below MMMU and the two are not comparable. Artificial Analysis publishes one MMMU-Pro figure rather than separate standard and vision splits, so this value should not be read as either split on its own.

77 models scored · top 10 shown

Ai2's precise instruction-following benchmark: 58 verifiable out-of-domain output constraints held out from the training constraints, so it measures generalisation to unseen instructions rather than familiarity with common ones. Scored as constraint-satisfaction accuracy.

137 models scored · top 10 shown

Professional chart-understanding benchmark with 100 real-world chart questions across 12 domains; preserve tool access and scoring setup because tool-enabled and no-tool results are distinct.

4 models scored · top 4 shown

Full informationsource

Chartography professional chart-understanding benchmark with tool access; keep separate from no-tool Chartography scores because tool use changes the evaluation setup.

1 model scored · top 1 shown

Math

MathArena contest-math leaderboard (AIME, AMC, olympiad-style); results depend on answer-format extraction and evaluation date.

4 models scored · top 4 shown

Full informationsource

Epoch AI's FrontierMath benchmark, Tier 4 (v2 problem set): exceptionally difficult, original research-level mathematics problems vetted by expert mathematicians. Tier 4 is not comparable to FrontierMath Tiers 1-3, and v2 problem sets are not comparable to v1 results.

5 models scored · top 5 shown

Full informationsource

Resolves real GitHub issues in popular Python repositories (500-task human-verified subset); agent scaffold, patch budget, and test policy materially affect the result.

39 models scored · top 10 shown

Long-horizon repository engineering across 1,865 tasks from 41 professional codebases; leaderboard results are agent-system scores, not model-only capability scores.

40 models scored · top 10 shown

Full informationsource

Real-world GitHub issue resolution across 300 tasks, 42 repositories, and 9 programming languages; agent scaffold and language mix matter.

21 models scored · top 10 shown

Contamination-resistant coding problems collected after model training cutoffs; score depends on the LiveCodeBench release and pass@k protocol.

94 models scored · top 10 shown

Full informationsource

Terminal-Bench 2.1 contains 89 agentic terminal tasks. Scores are agent-system results and vary with harness, effort level, timeout, resource allocation, task revision, and evaluation date; they are not model-only measurements.

110 models scored · top 10 shown

Full informationsource

Next-generation agentic terminal benchmark (Frontier-Bench); scores are agent-system results that vary with harness, effort, and evaluation date.

13 models scored · top 10 shown

Full informationsource

Semantically-versioned continuation of Terminal-Bench with recalibrated task resources and a refreshed, less-saturated task set; scores are agent-system results that vary with harness, effort level, and evaluation date, and are not directly comparable to Terminal-Bench 2.x or 3 results.

20 models scored · top 10 shown

Scientific research coding across physics, math, chemistry, biology, and materials — domain knowledge plus precise code synthesis; benchmark version and harness must match.

183 models scored · top 10 shown

Cursor's IDE-native agent eval on ambiguous, multi-file tasks from real Cursor sessions — correctness under Cursor's agent scaffolding (v3.2), not a model-only test.

20 models scored · top 10 shown

Contamination-resistant software engineering: rolling window of fresh GitHub issues under a fixed ReAct agent scaffold; scores are tied to the published task window.

64 models scored · top 10 shown

Long-horizon repository generation: build a full installable Python library from a natural-language spec, graded by upstream pytest; harness and task release are part of the result.

9 models scored · top 9 shown

DataCurve's long-horizon software-engineering benchmark of 113 original tasks across TypeScript, Go, Python, JavaScript, and Rust with program-based verifiers; scores are agent-system results under the eval harness.

29 models scored · top 10 shown

LMArena Code / WebDev human-preference Elo for front-end and agentic web development tasks; this is a time-varying leaderboard snapshot, not a fixed model property.

74 models scored · top 10 shown

Realistic function-level coding benchmark with 1,140 tasks across 7 languages; pass@1 depends on the release and prompting.

45 models scored · top 10 shown

Tool use & function calling

Realistic tool-agent user interactions in airline, retail, and banking domains; scores are agent-system results, not model-only capability.

3 models scored · top 3 shown

Full informationsource

Verified subset of Toolathlon, a tool-use benchmark measuring selecting, sequencing, and completing multi-step tasks with external tools; keep distinct from unverified Toolathlon runs.

27 models scored · top 10 shown

97 fintech customer-support tasks (disputes, account freezes, provisional credits, product changes) where an agent must navigate roughly 700 interconnected policy documents (~195k tokens) and execute multi-step tool calls. Graded pass@1 against backend database state rather than conversational quality. This is an agent-system result and is distinct from τ²-bench Telecom and from earlier τ-bench releases.

101 models scored · top 10 shown

Agents & computer use

Composite ranking →

Cybersecurity benchmark (UC Berkeley Sunblaze) evaluating AI agents on real-world vulnerability discovery and exploitation tasks; scores are agent-system results that vary with harness and environment.

8 models scored · top 8 shown

Full informationsource

Agent benchmark of long-horizon, real-world tasks requiring tool use and multi-step execution; results are agent-system measurements published by the evaluating provider.

10 models scored · top 10 shown

Full informationsource

Zapier benchmark evaluating how reliably agents complete realistic business workflows across 47 simulated SaaS tools; scores are agent-system results and depend on the scaffold.

16 models scored · top 10 shown

Full informationsource

Artificial Analysis's private agentic knowledge-work evaluation across four multi-week projects and 91 tasks. Its combined Elo metric includes rubric pass rate, analytical quality, and presentation quality; scores evaluate the agent system and its harness, not a model alone.

2 models scored · top 2 shown

Full informationsource

Agentic scientific-research workflows across the life, physical, Earth, mathematical, and engineering sciences, built on the Terminal-Bench harness with tasks authored by domain experts and verified programmatically. Scores are agent-system results with a wide per-model standard error (roughly ±3.5-4.5 pts) and are not comparable across benchmark releases.

15 models scored · top 10 shown

Artificial Analysis's agentic evaluation of OpenAI's GDPval task set, covering real-world knowledge work across dozens of occupations and industries. Scored via blind pairwise comparison fit to a Bradley-Terry Elo model anchored at 1000 for human-expert deliverables; requires shell and browser access through an agent harness (Stirrup), so it is not a raw model-only capability score.

24 models scored · top 10 shown

Full informationsource

OSWorld 2.0 long-horizon computer-use benchmark under partial-credit scoring, which awards credit for subtasks completed within a long real-world desktop workflow. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to pre-2.0 OSWorld results.

9 models scored · top 9 shown

Full informationsource

OSWorld 2.0 long-horizon computer-use benchmark under strict scoring, which requires the full task to complete for credit. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to partial-credit OSWorld 2.0 scores or to pre-2.0 OSWorld results.

3 models scored · top 3 shown

Full informationsource

OSWorld-V2 2.1 task release with partial-credit scoring. Keep separate from OSWorld 2.0 and preserve the agent harness, task release, and tool setup in each score's provenance.

1 model scored · top 1 shown

Arena

Blind pairwise human preference rating from the LMArena (Chatbot Arena) leaderboard; ratings drift with traffic, model aliases, sampling, and leaderboard methodology.

165 models scored · top 10 shown

Composite indices

Full informationsource

Artificial Analysis's own composite metric on a 0-100 scale, not a percentage. v4.2 is a weighted average of ten independently-run evaluations across four categories: agents 30% (AA-Briefcase 15%, GDPval-AA v2 10%, τ³-Banking 5%), coding 20% (Terminal-Bench 2.1 10%, SciCode 10%), general 30% (AA-Omniscience 15%, split 10% accuracy and 5% non-hallucination, GDP.pdf 10%, AA-LCR v1.1 5%), and scientific reasoning 20% (Humanity's Last Exam 10%, CritPt 10%). Index versions are not comparable to each other — v4.2 dropped GPQA Diamond and added AA-Briefcase and GDP.pdf relative to v4.1.1 — and several components are agent-harness evaluations, so the index is not a model-only capability measurement. Only listings Artificial Analysis has fully measured are recorded here; its estimated index values are omitted.

87 models scored · top 10 shown

Full informationsource

Artificial Analysis's own composite metric on a 0-100 scale, not a percentage. v4.3 is a weighted average of ten independently-run evaluations across four equally-weighted (25% each) categories: agents (AA-Briefcase, GDPval-AA v2, AutomationBench-AA), coding (Terminal-Bench 4.0, SciCode), general (AA-Omniscience, GDP.pdf, AA-LCR v1.1), and scientific reasoning (Humanity's Last Exam, CritPt). Relative to v4.2, v4.3 removes τ³-Banking, adds AutomationBench-AA, and swaps in Terminal-Bench 4.0 for Terminal-Bench 2.1. Index versions are not comparable to each other, so v4.3 scores must not be merged with v4.2 scores; several components are agent-harness evaluations, so the index is not a model-only capability measurement. Only listings Artificial Analysis has fully measured are recorded here; its estimated index values are omitted.

88 models scored · top 10 shown

Full informationsource

Artificial Analysis composite across AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. Keep separate from prior index versions because its evaluation components changed.

8 models scored · top 8 shown