LLMcompare

Search

Search for a command to run...

A · open · Gemma

Gemma 3 27B

Google

B · open · Gemma

Gemma 4 31B

Google

0:15

overall capability wins (1 ties across 16 scored). Gemma 4 31B leads overall.

  • Reasoning0:7
  • Coding0:4
  • Arena0:1
  • Tool use0:1
  • Composite indices0:2
  • Speed2:0
  • Specs0:2

Verdict

Gemma 4 31B leads coding (widest gap: +50.3 on LiveCode); Gemma 4 31B ranks higher on LMArena (+86 Elo); Gemma 4 31B offers ~2.0x larger context (256K vs 128K).

  • Gemma 4 31B leads coding (widest gap: +50.3 on LiveCode)
  • Gemma 4 31B ranks higher on LMArena (+86 Elo)
  • Gemma 4 31B offers ~2.0x larger context (256K vs 128K)

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Gemma 3 27B
$0.17
Gemma 4 31B
$0.17

Price per 1M tokens

Full comparison

Organization

Google

Google

Family

Gemma

Gemma

License

Gemma

Apache 2.0

Open weights

Yes

Yes

Release

Mar 12, 2025

Apr 8, 2026

Knowledge cutoff

-

-

API / provider

Google AI / self-host

Google AI / self-host

Modalities

text, image → text

text, image → text

Specs

Context window

128K

256K

B +128K

Max output

8K

8K

tie

Parameters

27B

31B

B +4.00

Pricing

Input $/1M

$0.10

$0.10

tie

Output $/1M

$0.30

$0.30

tie

Blended $/1M (3∶1)

$0.15

$0.15

tie

Speed

tok/s

90

85

A +5.00

TTFT (s)

0.28

0.3

A +0.02

Reasoning

MMLU-Pro

67.5

85.2

B +17.7 pts

GPQA Diamond

42.8

85.7

B +42.9 pts

Humanity's Last Exam

4.4

23.6

B +19.2 pts

AIME 2025

20.7

—

—

MATH-500

88.3

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

13.0

20.0

B +7.0 pts

AA-LCR v1.1

7.3

69.7

B +62.4 pts

CritPt

—

1.4

—

MMMU-Pro

48.0

73.4

B +25.4 pts

IFBench

31.8

75.6

B +43.8 pts

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

11.4

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

29.7

80.0

B +50.3 pts

Terminal-Bench 2.1

4.5

43.4

B +38.9 pts

Aider Polyglot

4.9

—

—

Terminal-Bench 3

—

—

—

BigCodeBench

—

—

—

SciCode

23.3

45.5

B +22.2 pts

CursorBench

—

—

—

SWE-Rebench

5.4

25.1

B +19.7 pts

NL2Repo-Bench

—

—

—

DeepSWE

—

—

—

WebDev Arena

—

1,364

—

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,365

1,451

B +86 Elo

Gemma 3 27B

Context
128K / 8K out
Parameters
27B
Price
$0.10 / $0.30
Speed
90 tok/s · 0.28s TTFT
Modalities
text, image → text
License
Gemma

Open multimodal Gemma 3 mid-size - strong preference Elo for its parameter count.

Gemma 4 31B

Context
256K / 8K out
Parameters
31B
Price
$0.10 / $0.30
Speed
85 tok/s · 0.3s TTFT
Modalities
text, image → text
License
Apache 2.0

Open-weight Gemma 4 instruction model - strong mid-size open alternative under Apache 2.0.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

MMLU-Pro: B +17.7GPQA: B +42.9HLE: B +19.2AA-LCR: B +62.4Omniscience Acc.: B +7.0MMMU-Pro: B +25.4IFBench: B +43.8

Coding

LiveCode: B +50.3TermBench: B +38.9SciCode: B +22.2SWE-Rebench: B +19.7

Tool use & function calling

τ³-Banking: B +14.0

Arena

Arena Elo: B +86

Composite indices

AA Index: B +10.5AA Index v4.3: B +10.5

Pick a different pair · Back to catalog