vs
0:15
overall capability wins (1 ties across 16 scored). Gemma 4 31B leads overall.
- Reasoning0:7
- Coding0:4
- Arena0:1
- Tool use0:1
- Composite indices0:2
- Speed2:0
- Specs0:2
Verdict
Gemma 4 31B leads coding (widest gap: +50.3 on LiveCode); Gemma 4 31B ranks higher on LMArena (+86 Elo); Gemma 4 31B offers ~2.0x larger context (256K vs 128K).
- Gemma 4 31B leads coding (widest gap: +50.3 on LiveCode)
- Gemma 4 31B ranks higher on LMArena (+86 Elo)
- Gemma 4 31B offers ~2.0x larger context (256K vs 128K)
Cost for 1M in + 250K out
Illustrative chat workload at primary-provider list prices.
- Gemma 3 27B
- $0.17
- Gemma 4 31B
- $0.17
Price per 1M tokens
Full comparison
Organization
Family
Gemma
Gemma
License
Gemma
Apache 2.0
Open weights
Yes
Yes
Release
Mar 12, 2025
Apr 8, 2026
Knowledge cutoff
-
-
API / provider
Google AI / self-host
Google AI / self-host
Modalities
text, image → text
text, image → text
Specs
Context window
128K
256K
Max output
8K
8K
Parameters
27B
31B
Pricing
Input $/1M
$0.10
$0.10
Output $/1M
$0.30
$0.30
Blended $/1M (3∶1)
$0.15
$0.15
Speed
tok/s
90
85
TTFT (s)
0.28
0.3
Reasoning
MMLU-Pro
67.5
85.2
GPQA Diamond
42.8
85.7
Humanity's Last Exam
4.4
23.6
AIME 2025
20.7
—
MATH-500
88.3
—
Humanity's Last Exam (with tools)
—
—
AA-Omniscience Accuracy
13.0
20.0
AA-LCR v1.1
7.3
69.7
CritPt
—
1.4
MMMU-Pro
48.0
73.4
IFBench
31.8
75.6
Chartography
—
—
Chartography (With Tools)
—
—
Coding
SWE-bench Verified
—
—
SWE-bench Pro
11.4
—
SWE-bench Multilingual
—
—
LiveCodeBench
29.7
80.0
Terminal-Bench 2.1
4.5
43.4
Aider Polyglot
4.9
—
Terminal-Bench 3
—
—
BigCodeBench
—
—
SciCode
23.3
45.5
CursorBench
—
—
SWE-Rebench
5.4
25.1
NL2Repo-Bench
—
—
DeepSWE
—
—
WebDev Arena
—
1,364
Terminal-Bench 4.0
—
—
CursorBench 4.0
—
—
FrontierCode v1.1 (Main)
—
—
Arena
LMArena Elo
1,365
1,451
| Metric | Gemma 3 27B | Gemma 4 31B | Delta |
|---|---|---|---|
| Identity | |||
| Organization | — | ||
| Family | Gemma | Gemma | — |
| License | Gemma | Apache 2.0 | — |
| Open weights | Yes | Yes | — |
| Release | Mar 12, 2025 | Apr 8, 2026 | — |
| Knowledge cutoff | - | - | — |
| API / provider | Google AI / self-host | Google AI / self-host | — |
| Modalities | text, image → text | text, image → text | — |
| Specs | |||
| Context window | 128K | 256K | B +128K |
| Max output | 8K | 8K | tie |
| Parameters | 27B | 31B | B +4.00 |
| Pricing | |||
| Input $/1M | $0.10 | $0.10 | tie |
| Output $/1M | $0.30 | $0.30 | tie |
| Blended $/1M (3∶1) | $0.15 | $0.15 | tie |
| Speed | |||
| tok/s | 90 | 85 | A +5.00 |
| TTFT (s) | 0.28 | 0.3 | A +0.02 |
| Reasoning | |||
| MMLU-Pro | 67.5 | 85.2 | B +17.7 pts |
| GPQA Diamond | 42.8 | 85.7 | B +42.9 pts |
| Humanity's Last Exam | 4.4 | 23.6 | B +19.2 pts |
| AIME 2025 | 20.7 | — | — |
| MATH-500 | 88.3 | — | — |
| Humanity's Last Exam (with tools) | — | — | — |
| AA-Omniscience Accuracy | 13.0 | 20.0 | B +7.0 pts |
| AA-LCR v1.1 | 7.3 | 69.7 | B +62.4 pts |
| CritPt | — | 1.4 | — |
| MMMU-Pro | 48.0 | 73.4 | B +25.4 pts |
| IFBench | 31.8 | 75.6 | B +43.8 pts |
| Chartography | — | — | — |
| Chartography (With Tools) | — | — | — |
| Coding | |||
| SWE-bench Verified | — | — | — |
| SWE-bench Pro | 11.4 | — | — |
| SWE-bench Multilingual | — | — | — |
| LiveCodeBench | 29.7 | 80.0 | B +50.3 pts |
| Terminal-Bench 2.1 | 4.5 | 43.4 | B +38.9 pts |
| Aider Polyglot | 4.9 | — | — |
| Terminal-Bench 3 | — | — | — |
| BigCodeBench | — | — | — |
| SciCode | 23.3 | 45.5 | B +22.2 pts |
| CursorBench | — | — | — |
| SWE-Rebench | 5.4 | 25.1 | B +19.7 pts |
| NL2Repo-Bench | — | — | — |
| DeepSWE | — | — | — |
| WebDev Arena | — | 1,364 | — |
| Terminal-Bench 4.0 | — | — | — |
| CursorBench 4.0 | — | — | — |
| FrontierCode v1.1 (Main) | — | — | — |
| Arena | |||
| LMArena Elo | 1,365 | 1,451 | B +86 Elo |
Gemma 3 27B
- Context
- 128K / 8K out
- Parameters
- 27B
- Price
- $0.10 / $0.30
- Speed
- 90 tok/s · 0.28s TTFT
- Modalities
- text, image → text
- License
- Gemma
Open multimodal Gemma 3 mid-size - strong preference Elo for its parameter count.
Gemma 4 31B
- Context
- 256K / 8K out
- Parameters
- 31B
- Price
- $0.10 / $0.30
- Speed
- 85 tok/s · 0.3s TTFT
- Modalities
- text, image → text
- License
- Apache 2.0
Open-weight Gemma 4 instruction model - strong mid-size open alternative under Apache 2.0.
Benchmark charts
Winner bars are emphasized. Per-benchmark deltas sit above each chart.