LLMcompare

Search

Search for a command to run...

Agents & computer use

CyberGym

Cybersecurity benchmark (UC Berkeley Sunblaze) evaluating AI agents on real-world vulnerability discovery and exploitation tasks; scores are agent-system results that vary with harness and environment.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Agents & computer use
Metric
Percent
Direction
Higher is better
Coverage
8 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Atria Dawn Preview

    Shanghai AI Labopen

    86.5%
  2. 286.2%
  3. 3
    GLM-5.3

    Zhipu AIopen

    84.5%
  4. 4
    Claude Mythos 5

    Anthropicclosed

    83.8%
  5. 5
    DeepSeek V4 Pro 0813

    DeepSeekclosed

    83.3%
  6. 676.7%
  7. 7
    DeepSeek V4 Pro

    DeepSeekopen

    52.7%
  8. 8
    DeepSeek V4 Flash

    DeepSeekopen

    38.7%

Back to all benchmarks