LLMcompare

Search

Search for a command to run...

A · closed · Claude Sonnet / Haiku

Claude Sonnet 5

Anthropic

B · closed · GPT-5.x

GPT-5.6 Terra

OpenAI

2:14

overall capability wins (1 ties across 17 scored). GPT-5.6 Terra leads overall.

  • Reasoning0:6
  • Coding1:5
  • Arena0:1
  • Agents1:0
  • Tool use0:1
  • Composite indices0:2
  • Pricing2:0
  • Speed0:2
  • Specs0:1

Verdict

GPT-5.6 Terra leads coding (widest gap: +9.1 on TermBench 4); GPT-5.6 Terra ranks higher on LMArena (+5 Elo); GPT-5.6 Terra offers larger context (1.1M vs 1M).

  • GPT-5.6 Terra leads coding (widest gap: +9.1 on TermBench 4)
  • GPT-5.6 Terra ranks higher on LMArena (+5 Elo)
  • GPT-5.6 Terra offers larger context (1.1M vs 1M)

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Claude Sonnet 5
$4.50
GPT-5.6 Terra
$5

Claude Sonnet 5 is about 1.1x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

Anthropic

OpenAI

Family

Claude Sonnet / Haiku

GPT-5.x

License

Proprietary

Proprietary

Open weights

No

No

Release

Jun 30, 2026

Jul 9, 2026

Knowledge cutoff

-

-

API / provider

Anthropic

OpenAI

Modalities

text, image → text

text, image → text

Specs

Context window

1M

1.1M

B +50K

Max output

128K

128K

tie

Parameters

—

—

—

Pricing

Input $/1M

$2

$2

tie

Output $/1M

$10

$12

A +$2

Blended $/1M (3∶1)

$4

$4.50

A +$0.50

Speed

tok/s

95

135

B +40

TTFT (s)

0.4

0.3

B +0.10

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

91.1

92.5

B +1.4 pts

Humanity's Last Exam

41.3

42.9

B +1.6 pts

AIME 2025

—

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

40.1

46.8

B +6.7 pts

AA-LCR v1.1

82.0

83.0

B +1.0 pts

CritPt

16.9

30.0

B +13.1 pts

MMMU-Pro

77.3

80.7

B +3.4 pts

IFBench

—

71.2

—

Chartography

15.6

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

63.4

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

80.5

88.0

B +7.5 pts

Aider Polyglot

—

—

—

Terminal-Bench 3

14.6

20.8

B +6.2 pts

BigCodeBench

—

—

—

SciCode

54.3

55.0

B +0.7 pts

CursorBench

61.5

64.9

B +3.4 pts

SWE-Rebench

56.8

—

—

NL2Repo-Bench

—

—

—

DeepSWE

54.0

—

—

WebDev Arena

1,537

1,521

A +16 Elo

Terminal-Bench 4.0

12.4

21.5

B +9.1 pts

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,461

1,466

B +5 Elo

Claude Sonnet 5

Context
1M / 128K out
Parameters
—
Price
$2 / $10
Speed
95 tok/s · 0.4s TTFT
Modalities
text, image → text
License
Proprietary

Default Claude model for Free/Pro at $2/$10 (the launch introductory rate, made permanent on 2026-08-10 - the planned 2026-09-01 increase to $3/$15 was cancelled); closes much of the gap to prior Opus on agentic tasks.

GPT-5.6 Terra

Context
1.1M / 128K out
Parameters
—
Price
$2 / $12
Speed
135 tok/s · 0.3s TTFT
Modalities
text, image → text
License
Proprietary

Balanced GPT-5.6 mid-tier for everyday work; July 30, 2026 API cut brought Standard rates to $2/$12 (20% lower), with strong throughput for the price.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: B +1.4HLE: B +1.6CritPt: B +13.1AA-LCR: B +1.0Omniscience Acc.: B +6.7MMMU-Pro: B +3.4

Coding

TermBench: B +7.5TermBench 3: B +6.2TermBench 4: B +9.1SciCode: B +0.7CursorBench: B +3.4WebDev Elo: A +16

Tool use & function calling

τ³-Banking: B +2.9

Agents & computer use

GDPval-AA: A +18

Arena

Arena Elo: B +5

Composite indices

AA Index: B +1.7AA Index v4.3: B +3.9

Pick a different pair · Back to catalog