LLMcompare

Search

Search for a command to run...

A · open · Mistral

Devstral 2

Mistral

B · closed · Mistral

Mistral Large 2

Mistral

7:0

overall capability wins (1 ties across 8 scored). Devstral 2 leads overall.

  • Reasoning4:0
  • Coding1:0
  • Pricing3:0
  • Speed2:0
  • Specs2:0

Verdict

Devstral 2 leads coding (widest gap: +3.6 on SciCode); Devstral 2 offers ~2.0x larger context (256K vs 128K); Devstral 2 allows ~7.8x more max output (64K vs 8K).

  • Devstral 2 leads coding (widest gap: +3.6 on SciCode)
  • Devstral 2 offers ~2.0x larger context (256K vs 128K)
  • Devstral 2 allows ~7.8x more max output (64K vs 8K)
  • Devstral 2 ships open weights (Apache 2.0)
  • Devstral 2 is ~3.8x cheaper on a blended token basis than Mistral Large 2

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Devstral 2
$0.90
Mistral Large 2
$3.50

Devstral 2 is about 3.9x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

Mistral

Mistral

Family

Mistral

Mistral

License

Apache 2.0

Mistral Research / Proprietary API

Open weights

Yes

No

Release

Dec 1, 2025

Jul 24, 2024

Knowledge cutoff

-

-

API / provider

Mistral

Mistral

Modalities

text → text

text → text

Specs

Context window

256K

128K

A +128K

Max output

64K

8K

A +56K

Parameters

123B

123B

tie

Pricing

Input $/1M

$0.40

$2

A +$1.60

Output $/1M

$2

$6

A +$4

Blended $/1M (3∶1)

$0.80

$3

A +$2.20

Speed

tok/s

85

65

A +20

TTFT (s)

0.35

0.45

A +0.10

Reasoning

MMLU-Pro

76.2

—

—

GPQA Diamond

59.4

48.6

A +10.8 pts

Humanity's Last Exam

3.6

3.3

A +0.3 pts

AIME 2025

36.7

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

20.8

19.9

A +0.9 pts

AA-LCR v1.1

32.3

—

—

CritPt

—

—

—

MMMU-Pro

—

—

—

IFBench

38.1

31.2

A +6.9 pts

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

44.8

—

—

Terminal-Bench 2.1

30.3

—

—

Aider Polyglot

—

—

—

Terminal-Bench 3

—

—

—

BigCodeBench

—

—

—

SciCode

32.8

29.2

A +3.6 pts

CursorBench

—

—

—

SWE-Rebench

42.0

—

—

NL2Repo-Bench

—

—

—

DeepSWE

—

—

—

WebDev Arena

1,194

—

—

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

—

—

—

Devstral 2

Context
256K / 64K out
Parameters
123B
Price
$0.40 / $2
Speed
85 tok/s · 0.35s TTFT
Modalities
text → text
License
Apache 2.0

Open-weight agentic coding model for autonomous software engineering (successor path: Medium 3.5).

Mistral Large 2

Context
128K / 8K out
Parameters
123B
Price
$2 / $6
Speed
65 tok/s · 0.45s TTFT
Modalities
text → text
License
Mistral Research / Proprietary API

123B flagship before Large 3 — EU-hosted frontier option with strong multilingual coding.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: A +10.8HLE: A +0.3Omniscience Acc.: A +0.9IFBench: A +6.9

Coding

SciCode: A +3.6

Pick a different pair · Back to catalog