LLMcompare

Search

Search for a command to run...

A · open · gpt-oss

gpt-oss-120B

OpenAI

B · open · gpt-oss

gpt-oss-20B

OpenAI

13:3

overall capability wins (1 ties across 17 scored). gpt-oss-120B leads overall.

  • Reasoning6:1
  • Coding3:2
  • Arena1:0
  • Tool use1:0
  • Composite indices2:0
  • Pricing0:3
  • Speed0:2
  • Specs2:0

Verdict

gpt-oss-120B leads coding (widest gap: +19.2 on SWE-Rebench); gpt-oss-120B ranks higher on LMArena (+35 Elo); gpt-oss-120B allows ~2.0x more max output (33K vs 16K).

  • gpt-oss-120B leads coding (widest gap: +19.2 on SWE-Rebench)
  • gpt-oss-120B ranks higher on LMArena (+35 Elo)
  • gpt-oss-120B allows ~2.0x more max output (33K vs 16K)
  • gpt-oss-20B is ~3.0x cheaper on a blended token basis than gpt-oss-120B

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

gpt-oss-120B
$0.30
gpt-oss-20B
$0.10

gpt-oss-20B is about 3.0x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

OpenAI

OpenAI

Family

gpt-oss

gpt-oss

License

Apache 2.0

Apache 2.0

Open weights

Yes

Yes

Release

Aug 5, 2025

Aug 5, 2025

Knowledge cutoff

-

-

API / provider

self-host / partners

self-host / partners

Modalities

text → text

text → text

Specs

Context window

131K

131K

tie

Max output

33K

16K

A +16K

Parameters

117B

21B

A +96

Pricing

Input $/1M

$0.15

$0.05

B +$0.10

Output $/1M

$0.60

$0.20

B +$0.40

Blended $/1M (3∶1)

$0.26

$0.09

B +$0.17

Speed

tok/s

70

120

B +50

TTFT (s)

0.4

0.2

B +0.20

Reasoning

MMLU-Pro

—

—

—

GPQA Diamond

78.2

68.8

A +9.4 pts

Humanity's Last Exam

19.6

11.0

A +8.6 pts

AIME 2025

92.5

91.7

A +0.8 pts

MATH-500

—

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

21.8

16.0

A +5.8 pts

AA-LCR v1.1

52.0

34.7

A +17.3 pts

CritPt

1.1

1.4

B +0.3 pts

MMMU-Pro

—

—

—

IFBench

69.0

65.1

A +3.9 pts

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

26.0

60.7

B +34.7 pts

SWE-bench Pro

16.2

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

—

—

—

Terminal-Bench 2.1

26.2

13.9

A +12.3 pts

Aider Polyglot

41.8

34.2

A +7.6 pts

Terminal-Bench 3

—

—

—

BigCodeBench

—

—

—

SciCode

34.0

38.9

B +4.9 pts

CursorBench

—

—

—

SWE-Rebench

27.3

8.1

A +19.2 pts

NL2Repo-Bench

—

—

—

DeepSWE

—

—

—

WebDev Arena

—

—

—

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,352

1,317

A +35 Elo

gpt-oss-120B

Context
131K / 33K out
Parameters
117B
Price
$0.15 / $0.60
Speed
70 tok/s · 0.4s TTFT
Modalities
text → text
License
Apache 2.0

OpenAI's open-weight 120B release under Apache 2.0 for self-hosted frontier-adjacent chat.

gpt-oss-20B

Context
131K / 16K out
Parameters
21B
Price
$0.05 / $0.20
Speed
120 tok/s · 0.2s TTFT
Modalities
text → text
License
Apache 2.0

Compact open gpt-oss variant that runs on a single consumer GPU.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

GPQA: A +9.4HLE: A +8.6AIME: A +0.8CritPt: B +0.3AA-LCR: A +17.3Omniscience Acc.: A +5.8IFBench: A +3.9

Coding

SWE-bench: B +34.7TermBench: A +12.3Aider: A +7.6SciCode: B +4.9SWE-Rebench: A +19.2

Tool use & function calling

τ³-Banking: A +5.8

Arena

Arena Elo: A +35

Composite indices

AA Index: A +6.6AA Index v4.3: A +3.3

Pick a different pair · Back to catalog