Agents & computer use
CyberGym
Cybersecurity benchmark (UC Berkeley Sunblaze) evaluating AI agents on real-world vulnerability discovery and exploitation tasks; scores are agent-system results that vary with harness and environment.
How models are tested
Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.
- Category
- Agents & computer use
- Metric
- Percent
- Direction
- Higher is better
- Coverage
- 8 / 318Models in this catalog with a published score
- Updated
- Catalog snapshot date
Scores
- 1Atria Dawn Preview86.5%
Shanghai AI Labopen
- 2Gemini 3.8 Flash Cyber86.2%
Googleclosed
- 3GLM-5.384.5%
Zhipu AIopen
- 4Claude Mythos 583.8%
Anthropicclosed
- 5DeepSeek V4 Pro 081383.3%
DeepSeekclosed
- 6DeepSeek V4 Flash 073176.7%
DeepSeekopen
- 7DeepSeek V4 Pro52.7%
DeepSeekopen
- 8DeepSeek V4 Flash38.7%
DeepSeekopen