An LLMcompare-computed index, not a published benchmark. Eligible benches are z-scored against models released in the past year. The index is a shrunken mean of those z-scores — sum(z) / (published count + k), with k dummy average results — so a new model that is far above peers on every measured bench is not treated as average on the rest, while a short brochure set still cannot dominate. Beating a long tail of older models does not inflate the index. Saturated and thinly scored benches are left out. Models with fewer than three published eligible scores are unranked, and every contributing score stays visible with its own provenance.
224 models scored · top 10 shown
- 1Claude Fable 5.1+1.01 · 16/25
Anthropicclosed
- 2GPT-6 Astra+0.88 · 13/25
OpenAIclosed
- 3Claude Opus 5+0.87 · 19/25
Anthropicclosed
- 4Claude Fable 5+0.74 · 20/25
Anthropicclosed
- 5Claude Opus 5.5+0.70 · 5/25
Anthropicclosed
- 6Muse Spark 1.3+0.69 · 9/25
Metaclosed
- 7GPT-5.6 Sol+0.68 · 21/25
OpenAIclosed
- 8Kimi K3+0.60 · 15/25
Moonshotopen
- 9Claude Mythos 5+0.57 · 8/25
Anthropicclosed
- 10Muse Spark 1.1+0.54 · 11/25
Metaclosed