LLMcompare

Search

Search for a command to run...

A · open · Phi

Phi-4

Microsoft

B · open · Phi

Phi-4 Reasoning

Microsoft

0:6

overall capability wins (1 ties across 7 scored). Phi-4 Reasoning leads overall.

  • Reasoning0:3
  • Coding0:1
  • Speed2:0
  • Specs0:2

Verdict

Phi-4 Reasoning leads coding (widest gap: +30.7 on LiveCode); Phi-4 Reasoning offers ~2.0x larger context (33K vs 16K); Phi-4 Reasoning allows ~4.0x more max output (16K vs 4K).

  • Phi-4 Reasoning leads coding (widest gap: +30.7 on LiveCode)
  • Phi-4 Reasoning offers ~2.0x larger context (33K vs 16K)
  • Phi-4 Reasoning allows ~4.0x more max output (16K vs 4K)

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Phi-4
$0.11
Phi-4 Reasoning
$0.11

Price per 1M tokens

Full comparison

Organization

Microsoft

Microsoft

Family

Phi

Phi

License

MIT

MIT

Open weights

Yes

Yes

Release

Dec 12, 2024

Apr 30, 2025

Knowledge cutoff

-

-

API / provider

Azure / self-host

Azure / self-host

Modalities

text → text

text → text

Specs

Context window

16K

33K

B +16K

Max output

4K

16K

B +12K

Parameters

14B

14B

tie

Pricing

Input $/1M

$0.07

$0.07

tie

Output $/1M

$0.14

$0.14

tie

Blended $/1M (3∶1)

$0.09

$0.09

tie

Speed

tok/s

150

80

A +70

TTFT (s)

0.15

0.5

A +0.35

Reasoning

MMLU-Pro

70.4

74.3

B +3.9 pts

GPQA Diamond

57.5

65.8

B +8.3 pts

Humanity's Last Exam

3.8

—

—

AIME 2025

18.0

62.9

B +44.9 pts

MATH-500

81.0

—

—

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

14.1

—

—

AA-LCR v1.1

—

—

—

CritPt

—

—

—

MMMU-Pro

—

—

—

IFBench

23.5

—

—

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

23.1

53.8

B +30.7 pts

Terminal-Bench 2.1

—

—

—

Aider Polyglot

—

—

—

Terminal-Bench 3

—

—

—

BigCodeBench

45.5

—

—

SciCode

26.0

—

—

CursorBench

—

—

—

SWE-Rebench

—

—

—

NL2Repo-Bench

—

—

—

DeepSWE

—

—

—

WebDev Arena

—

—

—

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,256

—

—

Phi-4

Context
16K / 4K out
Parameters
14B
Price
$0.07 / $0.14
Speed
150 tok/s · 0.15s TTFT
Modalities
text → text
License
MIT

Small high-quality dense model from Microsoft - punches above its size on STEM.

Phi-4 Reasoning

Context
33K / 16K out
Parameters
14B
Price
$0.07 / $0.14
Speed
80 tok/s · 0.5s TTFT
Modalities
text → text
License
MIT

Reasoning-tuned Phi-4 variant for math and science on constrained hardware.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

MMLU-Pro: B +3.9GPQA: B +8.3AIME: B +44.9

Coding

LiveCode: B +30.7

Pick a different pair · Back to catalog