Coding
CursorBench
Cursor's IDE-native agent eval on ambiguous, multi-file tasks from real Cursor sessions — correctness under Cursor's agent scaffolding (v3.2), not a model-only test.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Coding
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 20 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
- Source
- cursor.com/cursorbench
Scores
- 1Claude Fable 5.173.4%
Anthropicclosed
- 2Grok 4.670.8%
xAIclosed
- 3Claude Fable 570.5%
Anthropicclosed
- 4Claude Mythos 570.5%
Anthropicclosed
- 5Claude Opus 570.0%
Anthropicclosed
- 6Gemini 3.8 Flash69.2%
Googleclosed
- 7GPT-5.6 Sol67.2%
OpenAIclosed
- 8Cursor Grok 4.566.7%
Cursorclosed
- 9GPT-5.6 Terra64.9%
OpenAIclosed
- 10Claude Opus 4.862.3%
Anthropicclosed
- 11Gemini 3.7 Flash61.6%
Googleclosed
- 12Claude Sonnet 561.5%
Anthropicclosed
- 13GPT-5.6 Luna61.1%
OpenAIclosed
- 14Kimi K360.8%
Moonshotopen
- 15GPT-5.558.4%
OpenAIclosed
- 16Composer 2.556.1%
Cursorclosed
- 17GLM-5.255.0%
Zhipu AIopen
- 18Gemini 3.6 Flash53.5%
Googleclosed
- 19Kimi K2.7 Code49.7%
Moonshotopen
- 20Gemini 3.5 Flash48.8%
Googleclosed