LLMcompare

Search

Search for a command to run...

A · closed · Other

Mercury 2

Inception Labs

B · closed · Other

Step-2

StepFun

1:0

overall capability wins (0 ties across 1 scored). Mercury 2 leads overall.

  • Pricing3:0
  • Specs1:0

Verdict

Mercury 2 offers ~8.0x larger context (128K vs 16K); Mercury 2 is ~7.0x cheaper on a blended token basis than Step-2.

  • Mercury 2 offers ~8.0x larger context (128K vs 16K)
  • Mercury 2 is ~7.0x cheaper on a blended token basis than Step-2

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Mercury 2
$0.44
Step-2
$3

Mercury 2 is about 6.9x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

Inception Labs

StepFun

Family

Other

Other

License

Proprietary

Proprietary

Open weights

No

No

Release

Sep 1, 2026

Jan 15, 2025

Knowledge cutoff

-

-

API / provider

Inception Labs

StepFun

Modalities

text → text

text → text

Specs

Context window

128K

16K

A +112K

Max output

-

8K

—

Parameters

—

—

—

Pricing

Input $/1M

$0.25

$1.50

A +$1.25

Output $/1M

$0.75

$6

A +$5.25

Blended $/1M (3∶1)

$0.38

$2.63

A +$2.25

Speed

tok/s

-

50

—

TTFT (s)

-

0.6

—

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

77.0

—

—

Humanity's Last Exam

17.1

—

—

AIME 2025

—

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

21.2

—

—

AA-LCR v1.1

43.7

—

—

CritPt

0.8

—

—

MMMU-Pro

—

—

—

IFBench

69.8

—

—

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

27.3

—

—

Aider Polyglot

—

—

—

Terminal-Bench 3

—

—

—

BigCodeBench

—

—

—

SciCode

37.7

—

—

CursorBench

—

—

—

SWE-Rebench

—

—

—

NL2Repo-Bench

—

—

—

DeepSWE

—

—

—

WebDev Arena

1,167

—

—

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

—

—

—

Mercury 2

Context
128K
Parameters
—
Price
$0.25 / $0.75
Speed
—
Modalities
text → text
License
Proprietary

Inception Labs' diffusion-based (non-autoregressive) reasoning LLM, API-only with a 128K context window and high-throughput inference on Blackwell GPUs.

Step-2

Context
16K / 8K out
Parameters
—
Price
$1.50 / $6
Speed
50 tok/s · 0.6s TTFT
Modalities
text → text
License
Proprietary

StepFun frontier reasoning model that broke into Chinese and global arenas in early 2025.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

These two models have no overlapping published benchmarks in our dataset. Compare specs and pricing instead.

Pick a different pair · Back to catalog