LLMcompare

Search

Search for a command to run...

A · open · DeepSeek

DeepSeek V4 Pro

DeepSeek

B · open · Qwen

Qwen3.8 Max

Alibaba

2:14

overall capability wins (1 ties across 17 scored). Qwen3.8 Max leads overall.

  • Reasoning1:4
  • Coding0:5
  • Arena0:1
  • Agents0:1
  • Tool use0:2
  • Composite indices0:2
  • Pricing3:0
  • Speed0:1
  • Specs1:1

Verdict

Qwen3.8 Max leads coding (widest gap: +224.9 on WebDev Elo); Qwen3.8 Max ranks higher on LMArena (+23 Elo); DeepSeek V4 Pro allows ~2.9x more max output (384K vs 131K).

  • Qwen3.8 Max leads coding (widest gap: +224.9 on WebDev Elo)
  • Qwen3.8 Max ranks higher on LMArena (+23 Elo)
  • DeepSeek V4 Pro allows ~2.9x more max output (384K vs 131K)
  • DeepSeek V4 Pro is ~5.5x cheaper on a blended token basis than Qwen3.8 Max

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

DeepSeek V4 Pro
$0.65
Qwen3.8 Max
$3.50

DeepSeek V4 Pro is about 5.4x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

DeepSeek

Alibaba

Family

DeepSeek

Qwen

License

MIT

Qwen3.8-Max License

Open weights

Yes

Yes

Release

Apr 24, 2026

Aug 3, 2026

Knowledge cutoff

-

-

API / provider

DeepSeek

Alibaba Cloud

Modalities

text → text

text, image, video → text

Specs

Context window

1M

1M

tie

Max output

384K

131K

A +253K

Parameters

1.6T (49B act.)

2.4T (95B act.)

B +800

Pricing

Input $/1M

$0.43

$2

A +$1.56

Output $/1M

$0.87

$6

A +$5.13

Blended $/1M (3∶1)

$0.54

$3

A +$2.46

Speed

tok/s

85

86

B +1.00

TTFT (s)

0.4

0.4

tie

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

88.8

92.8

B +4.0 pts

Humanity's Last Exam

37.5

43.1

B +5.6 pts

AIME 2025

—

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

43.0

31.7

A +11.3 pts

AA-LCR v1.1

74.7

80.3

B +5.6 pts

CritPt

12.9

17.7

B +4.8 pts

MMMU-Pro

—

82.8

—

IFBench

76.5

—

—

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

67.7

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

64.0

81.3

B +17.3 pts

Aider Polyglot

—

—

—

Terminal-Bench 3

—

—

—

BigCodeBench

—

—

—

SciCode

50.8

52.1

B +1.3 pts

CursorBench

—

—

—

SWE-Rebench

41.4

—

—

NL2Repo-Bench

38.5

55.9

B +17.4 pts

DeepSWE

12.8

56.6

B +43.8 pts

WebDev Arena

1,446

1,671

B +225 Elo

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,457

1,481

B +23 Elo

DeepSeek V4 Pro

Context
1M / 384K out
Parameters
1.6T (49B act.)
Price
$0.43 / $0.87
Speed
85 tok/s · 0.4s TTFT
Modalities
text → text
License
MIT

Open-weight frontier MoE - exceptional quality-per-dollar with MIT license and 1M context. Scores shown are the original preview; the GA 0813 release improves agentic coding substantially.

Qwen3.8 Max

Context
1M / 131K out
Parameters
2.4T (95B act.)
Price
$2 / $6
Speed
86 tok/s · 0.4s TTFT
Modalities
text, image, video → text
License
Qwen3.8-Max License

Alibaba's Aug 2026 Max flagship (2.4T MoE / 95B active) — frontier agentic coding via QwenCloud / DashScope with 1M multimodal context and $2/$6 API pricing. Open weights released Aug 12, 2026 as Qwen3.8-2.4T-A95B (text-only, thinking-mandatory, Qwen3.8-Max license); the hosted version adds vision input, non-thinking mode, and built-in tools.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: B +4.0HLE: B +5.6CritPt: B +4.8AA-LCR: B +5.6Omniscience Acc.: A +11.3

Coding

TermBench: B +17.3SciCode: B +1.3NL2Repo: B +17.4DeepSWE: B +43.8WebDev Elo: B +225

Tool use & function calling

Toolathlon V.: B +16.6τ³-Banking: B +17.7

Agents & computer use

AutoBench: B +14.5

Arena

Arena Elo: B +23

Composite indices

AA Index: B +16.0AA Index v4.3: B +14.5

Pick a different pair · Back to catalog