LLMcompare

Search

Search for a command to run...

A · closed · Claude Sonnet / Haiku

Claude Sonnet 5

Anthropic

B · open · DeepSeek

DeepSeek V4 Pro

DeepSeek

12:2

overall capability wins (1 ties across 15 scored). Claude Sonnet 5 leads overall.

  • Reasoning4:1
  • Coding5:0
  • Arena1:0
  • Tool use2:0
  • Composite indices2:0
  • Pricing0:3
  • Speed1:0
  • Specs0:1

Verdict

Claude Sonnet 5 leads coding (widest gap: +91.0 on WebDev Elo); Claude Sonnet 5 ranks higher on LMArena (+4 Elo); DeepSeek V4 Pro allows ~3.0x more max output (384K vs 128K).

  • Claude Sonnet 5 leads coding (widest gap: +91.0 on WebDev Elo)
  • Claude Sonnet 5 ranks higher on LMArena (+4 Elo)
  • DeepSeek V4 Pro allows ~3.0x more max output (384K vs 128K)
  • DeepSeek V4 Pro ships open weights (MIT)
  • DeepSeek V4 Pro is ~7.4x cheaper on a blended token basis than Claude Sonnet 5

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Claude Sonnet 5
$4.50
DeepSeek V4 Pro
$0.65

DeepSeek V4 Pro is about 6.9x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

Anthropic

DeepSeek

Family

Claude Sonnet / Haiku

DeepSeek

License

Proprietary

MIT

Open weights

No

Yes

Release

Jun 30, 2026

Apr 24, 2026

Knowledge cutoff

-

-

API / provider

Anthropic

DeepSeek

Modalities

text, image → text

text → text

Specs

Context window

1M

1M

tie

Max output

128K

384K

B +256K

Parameters

—

1.6T (49B act.)

—

Pricing

Input $/1M

$2

$0.43

B +$1.56

Output $/1M

$10

$0.87

B +$9.13

Blended $/1M (3∶1)

$4

$0.54

B +$3.46

Speed

tok/s

95

85

A +10

TTFT (s)

0.4

0.4

tie

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

91.1

88.8

A +2.3 pts

Humanity's Last Exam

41.3

37.5

A +3.8 pts

AIME 2025

—

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

40.1

43.0

B +2.9 pts

AA-LCR v1.1

82.0

74.7

A +7.3 pts

CritPt

16.9

12.9

A +4.0 pts

MMMU-Pro

77.3

—

—

IFBench

—

76.5

—

Chartography

15.6

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

80.5

64.0

A +16.5 pts

Aider Polyglot

—

—

—

Terminal-Bench 3

14.6

—

—

BigCodeBench

—

—

—

SciCode

54.3

50.8

A +3.5 pts

CursorBench

61.5

—

—

SWE-Rebench

56.8

41.4

A +15.4 pts

NL2Repo-Bench

—

38.5

—

DeepSWE

54.0

12.8

A +41.2 pts

WebDev Arena

1,537

1,446

A +91 Elo

Terminal-Bench 4.0

12.4

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,461

1,457

A +4 Elo

Claude Sonnet 5

Context
1M / 128K out
Parameters
—
Price
$2 / $10
Speed
95 tok/s · 0.4s TTFT
Modalities
text, image → text
License
Proprietary

Default Claude model for Free/Pro at $2/$10 (the launch introductory rate, made permanent on 2026-08-10 - the planned 2026-09-01 increase to $3/$15 was cancelled); closes much of the gap to prior Opus on agentic tasks.

DeepSeek V4 Pro

Context
1M / 384K out
Parameters
1.6T (49B act.)
Price
$0.43 / $0.87
Speed
85 tok/s · 0.4s TTFT
Modalities
text → text
License
MIT

Open-weight frontier MoE - exceptional quality-per-dollar with MIT license and 1M context. Scores shown are the original preview; the GA 0813 release improves agentic coding substantially.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: A +2.3HLE: A +3.8CritPt: A +4.0AA-LCR: A +7.3Omniscience Acc.: B +2.9

Coding

TermBench: A +16.5SciCode: A +3.5SWE-Rebench: A +15.4DeepSWE: A +41.2WebDev Elo: A +91

Tool use & function calling

Toolathlon V.: A +15.7τ³-Banking: A +7.2

Arena

Arena Elo: A +4

Composite indices

AA Index: A +14.2AA Index v4.3: A +7.5

Pick a different pair · Back to catalog