LLMcompare

Search

Search for a command to run...

A · closed · Gemini

Gemini 3.6 Flash

Google

B · closed · GPT-5.x

GPT-5.6 Luna

OpenAI

6:9

overall capability wins (0 ties across 15 scored). GPT-5.6 Luna leads overall.

  • Reasoning4:2
  • Coding1:4
  • Arena1:0
  • Tool use0:1
  • Composite indices0:2
  • Pricing0:3
  • Speed0:2
  • Specs0:2

Verdict

GPT-5.6 Luna leads coding (widest gap: +20.0 on DeepSWE); Gemini 3.6 Flash edges reasoning & knowledge; Gemini 3.6 Flash ranks higher on LMArena (+27 Elo).

  • GPT-5.6 Luna leads coding (widest gap: +20.0 on DeepSWE)
  • Gemini 3.6 Flash edges reasoning & knowledge
  • Gemini 3.6 Flash ranks higher on LMArena (+27 Elo)
  • GPT-5.6 Luna offers larger context (1.1M vs 1M)
  • GPT-5.6 Luna allows ~2.0x more max output (128K vs 66K)
  • GPT-5.6 Luna is ~6.7x cheaper on a blended token basis than Gemini 3.6 Flash

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Gemini 3.6 Flash
$3.38
GPT-5.6 Luna
$0.50

GPT-5.6 Luna is about 6.8x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

Google

OpenAI

Family

Gemini

GPT-5.x

License

Proprietary

Proprietary

Open weights

No

No

Release

Jul 21, 2026

Jul 9, 2026

Knowledge cutoff

-

-

API / provider

Google

OpenAI

Modalities

text, image, audio, video → text

text, image → text

Specs

Context window

1M

1.1M

B +50K

Max output

66K

128K

B +62K

Parameters

—

—

—

Pricing

Input $/1M

$1.50

$0.20

B +$1.30

Output $/1M

$7.50

$1.20

B +$6.30

Blended $/1M (3∶1)

$3

$0.45

B +$2.55

Speed

tok/s

180

195

B +15

TTFT (s)

0.22

0.2

B +0.02

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

92.8

91.1

A +1.7 pts

Humanity's Last Exam

40.8

39.5

A +1.3 pts

AIME 2025

—

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

50.0

42.7

A +7.3 pts

AA-LCR v1.1

80.0

83.7

B +3.7 pts

CritPt

10.6

20.6

B +10.0 pts

MMMU-Pro

83.2

78.6

A +4.6 pts

IFBench

—

—

—

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

62.7

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

77.5

80.9

B +3.4 pts

Aider Polyglot

—

—

—

Terminal-Bench 3

—

14.3

—

BigCodeBench

—

—

—

SciCode

53.4

53.6

B +0.2 pts

CursorBench

53.5

61.1

B +7.6 pts

SWE-Rebench

—

43.6

—

NL2Repo-Bench

—

—

—

DeepSWE

47.0

67.0

B +20.0 pts

WebDev Arena

1,537

1,519

A +18 Elo

Terminal-Bench 4.0

—

17.3

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,480

1,453

A +27 Elo

Gemini 3.6 Flash

Context
1M / 66K out
Parameters
—
Price
$1.50 / $7.50
Speed
180 tok/s · 0.22s TTFT
Modalities
text, image, audio, video → text
License
Proprietary

Newest Flash model powering the free Gemini app - cheaper output tokens and strong price/performance.

GPT-5.6 Luna

Context
1.1M / 128K out
Parameters
—
Price
$0.20 / $1.20
Speed
195 tok/s · 0.2s TTFT
Modalities
text, image → text
License
Proprietary

Fastest, cheapest GPT-5.6 tier for high-volume chat, classification, and light coding; July 30, 2026 API cut dropped Standard rates 80% to $0.20/$1.20.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: A +1.7HLE: A +1.3CritPt: B +10.0AA-LCR: B +3.7Omniscience Acc.: A +7.3MMMU-Pro: A +4.6

Coding

TermBench: B +3.4SciCode: B +0.2CursorBench: B +7.6DeepSWE: B +20.0WebDev Elo: A +18

Tool use & function calling

τ³-Banking: B +1.2

Arena

Arena Elo: A +27

Composite indices

AA Index: B +3.1AA Index v4.3: B +3.2

Pick a different pair · Back to catalog