LLMcompare

Search

Search for a command to run...

A · closed · Gemini

Gemini 2.5 Pro

Google

B · closed · GPT-4 / o-series

GPT-4.1

OpenAI

15:1

overall capability wins (1 ties across 17 scored). Gemini 2.5 Pro leads overall.

  • Reasoning9:0
  • Coding4:1
  • Arena1:0
  • Pricing2:1
  • Speed1:1
  • Specs1:0

Verdict

Gemini 2.5 Pro leads coding (widest gap: +34.4 on LiveCode); Gemini 2.5 Pro ranks higher on LMArena (+31 Elo); Gemini 2.5 Pro allows ~2.0x more max output (66K vs 33K).

  • Gemini 2.5 Pro leads coding (widest gap: +34.4 on LiveCode)
  • Gemini 2.5 Pro ranks higher on LMArena (+31 Elo)
  • Gemini 2.5 Pro allows ~2.0x more max output (66K vs 33K)

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Gemini 2.5 Pro
$3.75
GPT-4.1
$4

Gemini 2.5 Pro is about 1.1x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

Google

OpenAI

Family

Gemini

GPT-4 / o-series

License

Proprietary

Proprietary

Open weights

No

No

Release

Mar 25, 2025

Apr 14, 2025

Knowledge cutoff

-

2024-06

API / provider

Google

OpenAI

Modalities

text, image, audio, video → text

text, image → text

Specs

Context window

1M

1M

tie

Max output

66K

33K

A +33K

Parameters

—

—

—

Pricing

Input $/1M

$1.25

$2

A +$0.75

Output $/1M

$10

$8

B +$2

Blended $/1M (3∶1)

$3.44

$3.50

A +$0.06

Speed

tok/s

100

95

A +5.00

TTFT (s)

0.45

0.35

B +0.10

Reasoning

MMLU-Pro

86.2

80.6

A +5.6 pts

GPQA Diamond

84.4

66.6

A +17.8 pts

Humanity's Last Exam

22.5

4.2

A +18.3 pts

AIME 2025

87.7

46.4

A +41.3 pts

MATH-500

96.7

91.3

A +5.4 pts

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

39.1

27.8

A +11.3 pts

AA-LCR v1.1

69.0

68.3

A +0.7 pts

CritPt

2.6

—

—

MMMU-Pro

74.9

61.2

A +13.7 pts

IFBench

48.7

43.0

A +5.7 pts

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

53.6

39.6

A +14.0 pts

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

80.1

45.7

A +34.4 pts

Terminal-Bench 2.1

28.5

—

—

Aider Polyglot

79.1

52.4

A +26.7 pts

Terminal-Bench 3

—

—

—

BigCodeBench

—

—

—

SciCode

46.3

38.1

A +8.2 pts

CursorBench

—

—

—

SWE-Rebench

23.5

30.1

B +6.6 pts

NL2Repo-Bench

—

—

—

DeepSWE

—

—

—

WebDev Arena

1,226

—

—

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,446

1,415

A +31 Elo

Gemini 2.5 Pro

Context
1M / 66K out
Parameters
—
Price
$1.25 / $10
Speed
100 tok/s · 0.45s TTFT
Modalities
text, image, audio, video → text
License
Proprietary

Previous-gen Google Pro model; still a capable long-context multimodal option at competitive rates.

GPT-4.1

Context
1M / 33K out
Parameters
—
Price
$2 / $8
Speed
95 tok/s · 0.35s TTFT
Modalities
text, image → text
License
Proprietary

Previous-gen OpenAI instruction model with 1M context; still useful for long-document workloads.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

MMLU-Pro: A +5.6GPQA: A +17.8HLE: A +18.3AIME: A +41.3MATH-500: A +5.4AA-LCR: A +0.7Omniscience Acc.: A +11.3MMMU-Pro: A +13.7IFBench: A +5.7

Coding

SWE-bench: A +14.0LiveCode: A +34.4Aider: A +26.7SciCode: A +8.2SWE-Rebench: B +6.6

Arena

Arena Elo: A +31

Pick a different pair · Back to catalog