LLMcompare

Search

Search for a command to run...

Reasoning & knowledge

Chartography

Professional chart-understanding benchmark with 100 real-world chart questions across 12 domains; preserve tool access and scoring setup because tool-enabled and no-tool results are distinct.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Reasoning & knowledge
Metric
Percent
Direction
Higher is better
Coverage
4 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Opus 5.5

    Anthropicclosed

    64.4%
  2. 2
    Claude Sonnet 5.5

    Anthropicclosed

    61.6%
  3. 3
    GPT-6 Sol

    OpenAIclosed

    53.6%
  4. 4
    Claude Sonnet 5

    Anthropicclosed

    15.6%

Back to all benchmarks