LLMcompare

Search

Search for a command to run...

A · closed · Kimi

Kimi K2.6

Moonshot

B · open · Kimi

Kimi K3

Moonshot

0:14

overall capability wins (0 ties across 14 scored). Kimi K3 leads overall.

  • Reasoning0:6
  • Coding0:3
  • Arena0:1
  • Tool use0:2
  • Composite indices0:2
  • Pricing3:0
  • Speed2:0
  • Specs0:2

Verdict

Kimi K3 leads coding (widest gap: +165.5 on WebDev Elo); Kimi K3 ranks higher on LMArena (+24 Elo); Kimi K3 offers ~4.1x larger context (1.0M vs 256K).

  • Kimi K3 leads coding (widest gap: +165.5 on WebDev Elo)
  • Kimi K3 ranks higher on LMArena (+24 Elo)
  • Kimi K3 offers ~4.1x larger context (1.0M vs 256K)
  • Kimi K3 allows ~3.9x more max output (64K vs 16K)
  • Kimi K3 ships open weights (Kimi K3 License)
  • Kimi K2.6 is ~6.9x cheaper on a blended token basis than Kimi K3

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Kimi K2.6
$1
Kimi K3
$6.75

Kimi K2.6 is about 6.8x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

Moonshot

Moonshot

Family

Kimi

Kimi

License

Proprietary

Kimi K3 License

Open weights

No

Yes

Release

Apr 8, 2026

Jul 16, 2026

Knowledge cutoff

-

-

API / provider

Moonshot

Moonshot

Modalities

text, image → text

text, image, video → text

Specs

Context window

256K

1.0M

B +793K

Max output

16K

64K

B +48K

Parameters

—

2.8T (104B act.)

—

Pricing

Input $/1M

$0.50

$3

A +$2.50

Output $/1M

$2

$15

A +$13

Blended $/1M (3∶1)

$0.88

$6

A +$5.13

Speed

tok/s

70

60

A +10

TTFT (s)

0.4

0.6

A +0.20

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

91.1

93.5

B +2.4 pts

Humanity's Last Exam

37.5

46.9

B +9.4 pts

AIME 2025

—

—

—

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

32.6

47.6

B +15.0 pts

AA-LCR v1.1

81.0

88.7

B +7.7 pts

CritPt

8.0

23.4

B +15.4 pts

MMMU-Pro

79.4

80.5

B +1.1 pts

IFBench

76.0

—

—

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

65.9

85.0

B +19.1 pts

Aider Polyglot

—

—

—

Terminal-Bench 3

—

—

—

BigCodeBench

—

—

—

SciCode

51.5

59.5

B +8.0 pts

CursorBench

—

60.8

—

SWE-Rebench

46.5

—

—

NL2Repo-Bench

—

—

—

DeepSWE

—

69.0

—

WebDev Arena

1,509

1,674

B +166 Elo

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,460

1,485

B +24 Elo

Kimi K2.6

Context
256K / 16K out
Parameters
—
Price
$0.50 / $2
Speed
70 tok/s · 0.4s TTFT
Modalities
text, image → text
License
Proprietary

Post-K2.5 Moonshot API SKU with longer context and stronger multimodal agent tooling.

Kimi K3

Context
1.0M / 64K out
Parameters
2.8T (104B act.)
Price
$3 / $15
Speed
60 tok/s · 0.6s TTFT
Modalities
text, image, video → text
License
Kimi K3 License

Largest open-weight model (2.8T MoE / 104B active) — official card reports strong reasoning and agentic coding results.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: B +2.4HLE: B +9.4CritPt: B +15.4AA-LCR: B +7.7Omniscience Acc.: B +15.0MMMU-Pro: B +1.1

Coding

TermBench: B +19.1SciCode: B +8.0WebDev Elo: B +166

Tool use & function calling

Toolathlon V.: B +18.5τ³-Banking: B +22.7

Arena

Arena Elo: B +24

Composite indices

AA Index: B +22.7AA Index v4.3: B +16.3

Pick a different pair · Back to catalog