LLMcompare

Search

Search for a command to run...

A · closed · Other

GPT-3.5 Turbo

OpenAI

B · closed · Other

Mercury 2

Inception Labs

0:2

overall capability wins (0 ties across 2 scored). Mercury 2 leads overall.

  • Reasoning0:1
  • Pricing0:3
  • Specs0:1

Verdict

Mercury 2 edges reasoning & knowledge; Mercury 2 offers ~7.8x larger context (128K vs 16K); Mercury 2 is ~2.0x cheaper on a blended token basis than GPT-3.5 Turbo.

  • Mercury 2 edges reasoning & knowledge
  • Mercury 2 offers ~7.8x larger context (128K vs 16K)
  • Mercury 2 is ~2.0x cheaper on a blended token basis than GPT-3.5 Turbo

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

GPT-3.5 Turbo
$0.88
Mercury 2
$0.44

Mercury 2 is about 2.0x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

OpenAI

Inception Labs

Family

Other

Other

License

Proprietary

Proprietary

Open weights

No

No

Release

Mar 1, 2023

Sep 1, 2026

Knowledge cutoff

-

-

API / provider

OpenAI

Inception Labs

Modalities

text → text

text → text

Specs

Context window

16K

128K

B +112K

Max output

4K

-

—

Parameters

—

—

—

Pricing

Input $/1M

$0.50

$0.25

B +$0.25

Output $/1M

$1.50

$0.75

B +$0.75

Blended $/1M (3∶1)

$0.75

$0.38

B +$0.38

Speed

tok/s

120

-

—

TTFT (s)

0.25

-

—

Reasoning

MMLU-Pro

46.2

—

—

GPQA Diamond

29.7

77.0

B +47.3 pts

Humanity's Last Exam

—

17.1

—

AIME 2025

—

—

—

MATH-500

44.1

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

—

21.2

—

AA-LCR v1.1

—

43.7

—

CritPt

—

0.8

—

MMMU-Pro

—

—

—

IFBench

—

69.8

—

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

—

27.3

—

Aider Polyglot

—

—

—

Terminal-Bench 3

—

—

—

BigCodeBench

—

—

—

SciCode

—

37.7

—

CursorBench

—

—

—

SWE-Rebench

—

—

—

NL2Repo-Bench

—

—

—

DeepSWE

—

—

—

WebDev Arena

—

1,167

—

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,225

—

—

GPT-3.5 Turbo

Context
16K / 4K out
Parameters
—
Price
$0.50 / $1.50
Speed
120 tok/s · 0.25s TTFT
Modalities
text → text
License
Proprietary

Classic ChatGPT-era budget model — still referenced for cheap classification and simple chat.

Mercury 2

Context
128K
Parameters
—
Price
$0.25 / $0.75
Speed
—
Modalities
text → text
License
Proprietary

Inception Labs' diffusion-based (non-autoregressive) reasoning LLM, API-only with a 128K context window and high-throughput inference on Blackwell GPUs.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: B +47.3

Pick a different pair · Back to catalog