LLMcompare

Search

Search for a command to run...

A · open · DeepSeek

DeepSeek V4 Pro

DeepSeek

B · open · Llama / Muse

Llama 4 Maverick

Meta

12:0

overall capability wins (1 ties across 13 scored). DeepSeek V4 Pro leads overall.

  • Reasoning5:0
  • Coding3:0
  • Arena1:0
  • Tool use1:0
  • Composite indices2:0
  • Pricing0:3
  • Speed0:2
  • Specs2:0

Verdict

DeepSeek V4 Pro leads coding (widest gap: +56.1 on TermBench); DeepSeek V4 Pro ranks higher on LMArena (+130 Elo); DeepSeek V4 Pro allows ~11.7x more max output (384K vs 33K).

  • DeepSeek V4 Pro leads coding (widest gap: +56.1 on TermBench)
  • DeepSeek V4 Pro ranks higher on LMArena (+130 Elo)
  • DeepSeek V4 Pro allows ~11.7x more max output (384K vs 33K)
  • Llama 4 Maverick is ~1.3x cheaper on a blended token basis than DeepSeek V4 Pro

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

DeepSeek V4 Pro
$0.65
Llama 4 Maverick
$0.48

Llama 4 Maverick is about 1.4x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

DeepSeek

Meta

Family

DeepSeek

Llama / Muse

License

MIT

Llama 4 Community

Open weights

Yes

Yes

Release

Apr 24, 2026

Apr 5, 2025

Knowledge cutoff

-

-

API / provider

DeepSeek

Together / Fireworks (ref.)

Modalities

text → text

text, image → text

Specs

Context window

1M

1M

tie

Max output

384K

33K

A +351K

Parameters

1.6T (49B act.)

400B (17B act.)

A +1K

Pricing

Input $/1M

$0.43

$0.27

B +$0.16

Output $/1M

$0.87

$0.85

B +$0.02

Blended $/1M (3∶1)

$0.54

$0.42

B +$0.13

Speed

tok/s

85

95

B +10

TTFT (s)

0.4

0.35

B +0.05

Reasoning

MMLU-Pro

—

80.5

—

GPQA Diamond

88.8

67.1

A +21.7 pts

Humanity's Last Exam

37.5

4.9

A +32.6 pts

AIME 2025

—

19.3

—

MATH-500

—

88.9

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

43.0

24.9

A +18.1 pts

AA-LCR v1.1

74.7

50.0

A +24.7 pts

CritPt

12.9

—

—

MMMU-Pro

—

62.1

—

IFBench

76.5

43.0

A +33.5 pts

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

21.0

—

SWE-bench Pro

—

5.2

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

43.4

—

Terminal-Bench 2.1

64.0

7.9

A +56.1 pts

Aider Polyglot

—

15.6

—

Terminal-Bench 3

—

—

—

BigCodeBench

—

49.7

—

SciCode

50.8

31.7

A +19.1 pts

CursorBench

—

—

—

SWE-Rebench

41.4

11.0

A +30.4 pts

NL2Repo-Bench

38.5

—

—

DeepSWE

12.8

—

—

WebDev Arena

1,446

—

—

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,457

1,327

A +130 Elo

DeepSeek V4 Pro

Context
1M / 384K out
Parameters
1.6T (49B act.)
Price
$0.43 / $0.87
Speed
85 tok/s · 0.4s TTFT
Modalities
text → text
License
MIT

Open-weight frontier MoE - exceptional quality-per-dollar with MIT license and 1M context. Scores shown are the original preview; the GA 0813 release improves agentic coding substantially.

Llama 4 Maverick

Context
1M / 33K out
Parameters
400B (17B act.)
Price
$0.27 / $0.85
Speed
95 tok/s · 0.35s TTFT
Modalities
text, image → text
License
Llama 4 Community

Meta's open multimodal MoE flagship - 400B total / 17B active with 1M context.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: A +21.7HLE: A +32.6AA-LCR: A +24.7Omniscience Acc.: A +18.1IFBench: A +33.5

Coding

TermBench: A +56.1SciCode: A +19.1SWE-Rebench: A +30.4

Tool use & function calling

τ³-Banking: A +26.4

Arena

Arena Elo: A +130

Composite indices

AA Index: A +21.6AA Index v4.3: A +21.6

Pick a different pair · Back to catalog