LLMcompare

Search

Search for a command to run...

A · closed · Claude Opus

Claude Opus 4.7

Anthropic

B · closed · Claude Opus

Claude Opus 4.8

Anthropic

6:8

overall capability wins (2 ties across 16 scored). Claude Opus 4.8 leads overall.

  • Reasoning2:4
  • Coding2:3
  • Arena1:0
  • Agents0:1
  • Tool use1:0
  • Speed0:1

Verdict

Claude Opus 4.8 leads coding (widest gap: +4.9 on SWE-Pro); Claude Opus 4.7 ranks higher on LMArena (+21 Elo).

  • Claude Opus 4.8 leads coding (widest gap: +4.9 on SWE-Pro)
  • Claude Opus 4.7 ranks higher on LMArena (+21 Elo)

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Claude Opus 4.7
$11.25
Claude Opus 4.8
$11.25

Price per 1M tokens

Full comparison

Organization

Anthropic

Anthropic

Family

Claude Opus

Claude Opus

License

Proprietary

Proprietary

Open weights

No

No

Release

Apr 16, 2026

May 28, 2026

Knowledge cutoff

-

-

API / provider

Anthropic

Anthropic

Modalities

text, image → text

text, image → text

Specs

Context window

1M

1M

tie

Max output

128K

128K

tie

Parameters

—

—

—

Pricing

Input $/1M

$5

$5

tie

Output $/1M

$25

$25

tie

Blended $/1M (3∶1)

$10

$10

tie

Speed

tok/s

68

72

B +4.00

TTFT (s)

0.55

0.55

tie

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

91.4

92.0

B +0.6 pts

Humanity's Last Exam

42.3

48.7

B +6.4 pts

AIME 2025

—

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

48.9

48.8

A +0.1 pts

AA-LCR v1.1

78.7

77.7

A +1.0 pts

CritPt

12.0

20.9

B +8.9 pts

MMMU-Pro

78.8

—

—

IFBench

58.6

62.2

B +3.6 pts

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

64.3

69.2

B +4.9 pts

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

83.1

84.6

B +1.5 pts

Aider Polyglot

—

—

—

Terminal-Bench 3

—

21.1

—

BigCodeBench

—

—

—

SciCode

54.5

54.4

A +0.1 pts

CursorBench

—

62.3

—

SWE-Rebench

53.1

56.5

B +3.4 pts

NL2Repo-Bench

—

—

—

DeepSWE

—

59.0

—

WebDev Arena

1,557

1,539

A +19 Elo

Terminal-Bench 4.0

—

23.6

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,494

1,473

A +21 Elo

Claude Opus 4.7

Context
1M / 128K out
Parameters
—
Price
$5 / $25
Speed
68 tok/s · 0.55s TTFT
Modalities
text, image → text
License
Proprietary

April 2026 Opus release with a large SWE-bench jump; still widely used alongside Opus 4.8 / 5.

Claude Opus 4.8

Context
1M / 128K out
Parameters
—
Price
$5 / $25
Speed
72 tok/s · 0.55s TTFT
Modalities
text, image → text
License
Proprietary

Previous Opus flagship with 1M context; still competitive for coding agents and computer use.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: B +0.6HLE: B +6.4CritPt: B +8.9AA-LCR: A +1.0Omniscience Acc.: A +0.1IFBench: B +3.6

Coding

SWE-Pro: B +4.9TermBench: B +1.5SciCode: A +0.1SWE-Rebench: B +3.4WebDev Elo: A +19

Tool use & function calling

τ³-Banking: A +0.4

Agents & computer use

GDPval-AA: B +95

Arena

Arena Elo: A +21

Pick a different pair · Back to catalog