LLMcompare

Search

Search for a command to run...

A · closed · Other

Mercury 2

Inception Labs

B · open · Other

Trinity Large Thinking

Arcee AI

6:5

overall capability wins (0 ties across 11 scored). Mercury 2 leads overall.

  • Reasoning4:2
  • Coding1:2
  • Tool use1:0
  • Composite indices2:0
  • Specs0:1

Verdict

Trinity Large Thinking leads coding (widest gap: +70.0 on WebDev Elo); Mercury 2 edges reasoning & knowledge; Trinity Large Thinking offers ~2.0x larger context (262K vs 128K).

  • Trinity Large Thinking leads coding (widest gap: +70.0 on WebDev Elo)
  • Mercury 2 edges reasoning & knowledge
  • Trinity Large Thinking offers ~2.0x larger context (262K vs 128K)
  • Trinity Large Thinking ships open weights (Apache 2.0)

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Mercury 2
$0.44
Trinity Large Thinking
-

Price per 1M tokens

Full comparison

Organization

Inception Labs

Arcee AI

Family

Other

Other

License

Proprietary

Apache 2.0

Open weights

No

Yes

Release

Sep 1, 2026

Jan 27, 2026

Knowledge cutoff

-

-

API / provider

Inception Labs

No primary API price listed

Modalities

text → text

text → text

Specs

Context window

128K

262K

B +134K

Max output

-

-

—

Parameters

—

400B (13B act.)

—

Pricing

Input $/1M

$0.25

—

—

Output $/1M

$0.75

—

—

Blended $/1M (3∶1)

$0.38

-

—

Speed

tok/s

-

-

—

TTFT (s)

-

-

—

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

77.0

75.2

A +1.8 pts

Humanity's Last Exam

17.1

15.8

A +1.3 pts

AIME 2025

—

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

21.2

22.5

B +1.3 pts

AA-LCR v1.1

43.7

38.0

A +5.7 pts

CritPt

0.8

0.9

B +0.1 pts

MMMU-Pro

—

—

—

IFBench

69.8

56.3

A +13.5 pts

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

27.3

20.6

A +6.7 pts

Aider Polyglot

—

—

—

Terminal-Bench 3

—

—

—

BigCodeBench

—

—

—

SciCode

37.7

40.6

B +2.9 pts

CursorBench

—

—

—

SWE-Rebench

—

—

—

NL2Repo-Bench

—

—

—

DeepSWE

—

—

—

WebDev Arena

1,167

1,237

B +70 Elo

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

—

1,369

—

Mercury 2

Context
128K
Parameters
—
Price
$0.25 / $0.75
Speed
—
Modalities
text → text
License
Proprietary

Inception Labs' diffusion-based (non-autoregressive) reasoning LLM, API-only with a 128K context window and high-throughput inference on Blackwell GPUs.

Trinity Large Thinking

Context
262K
Parameters
400B (13B act.)
Price
—
Speed
—
Modalities
text → text
License
Apache 2.0

Arcee AI's reasoning-tuned Trinity Large (400B total / 13B active MoE, 4-of-256 experts), Apache 2.0 open weights, larger sibling to Trinity Mini; released Jan 27, 2026.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: A +1.8HLE: A +1.3CritPt: B +0.1AA-LCR: A +5.7Omniscience Acc.: B +1.3IFBench: A +13.5

Coding

TermBench: A +6.7SciCode: B +2.9WebDev Elo: B +70

Tool use & function calling

τ³-Banking: A +3.7

Composite indices

AA Index: A +0.6AA Index v4.3: A +0.6

Pick a different pair · Back to catalog