LLMcompare

Search

Search for a command to run...

A · closed · Grok

Grok 4.5

xAI

B · closed · Grok

Grok 4.6

xAI

2:11

overall capability wins (1 ties across 14 scored). Grok 4.6 leads overall.

  • Reasoning1:4
  • Coding0:5
  • Arena1:0
  • Agents0:1
  • Tool use0:1
  • Composite indices0:2

Verdict

Grok 4.6 leads coding (widest gap: +62.5 on WebDev Elo); Grok 4.5 ranks higher on LMArena (+13 Elo).

  • Grok 4.6 leads coding (widest gap: +62.5 on WebDev Elo)
  • Grok 4.5 ranks higher on LMArena (+13 Elo)

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Grok 4.5
$3.50
Grok 4.6
$3.50

Price per 1M tokens

Full comparison

Organization

xAI

xAI

Family

Grok

Grok

License

Proprietary

Proprietary

Open weights

No

No

Release

Jul 9, 2026

Aug 12, 2026

Knowledge cutoff

-

2026-02-01

API / provider

xAI

xAI

Modalities

text, image → text

text, image → text

Specs

Context window

500K

500K

tie

Max output

128K

-

—

Parameters

—

1.5T

—

Pricing

Input $/1M

$2

$2

tie

Output $/1M

$6

$6

tie

Blended $/1M (3∶1)

$3

$3

tie

Speed

tok/s

100

-

—

TTFT (s)

0.35

-

—

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

93.1

94.9

B +1.8 pts

Humanity's Last Exam

42.7

42.9

B +0.2 pts

AIME 2025

—

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

51.6

48.2

A +3.4 pts

AA-LCR v1.1

79.3

80.3

B +1.0 pts

CritPt

15.4

17.1

B +1.7 pts

MMMU-Pro

80.4

—

—

IFBench

—

—

—

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

81.6

88.4

B +6.8 pts

Aider Polyglot

—

—

—

Terminal-Bench 3

15.7

26.5

B +10.8 pts

BigCodeBench

—

—

—

SciCode

55.0

56.5

B +1.5 pts

CursorBench

—

70.8

—

SWE-Rebench

63.8

—

—

NL2Repo-Bench

—

—

—

DeepSWE

—

67.0

—

WebDev Arena

1,555

1,618

B +63 Elo

Terminal-Bench 4.0

12.4

20.3

B +7.9 pts

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,468

1,456

A +13 Elo

Grok 4.5

Context
500K / 128K out
Parameters
—
Price
$2 / $6
Speed
100 tok/s · 0.35s TTFT
Modalities
text, image → text
License
Proprietary

xAI/SpaceXAI public API flagship (July 2026) - Cursor-trained coding focus at $2/$6. For the Cursor IDE jointly-trained SKU and Fast tier, see Cursor Grok 4.5.

Grok 4.6

Context
500K
Parameters
1.5T
Price
$2 / $6
Speed
—
Modalities
text, image → text
License
Proprietary

xAI/SpaceXAI's August 2026 flagship (1.5T-parameter) tuned for long-running agents and ambitious interactive/visual work — 500k context at $2/$6, matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index (61). A fast variant is available at 2x price.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: B +1.8HLE: B +0.2CritPt: B +1.7AA-LCR: B +1.0Omniscience Acc.: A +3.4

Coding

TermBench: B +6.8TermBench 3: B +10.8TermBench 4: B +7.9SciCode: B +1.5WebDev Elo: B +63

Tool use & function calling

τ³-Banking: B +8.6

Agents & computer use

GDPval-AA: B +237

Arena

Arena Elo: A +13

Composite indices

AA Index: B +5.1AA Index v4.3: B +5.3

Pick a different pair · Back to catalog