LLMcompare

Search

Search for a command to run...

A · open · Phi

Phi-3 Mini 3.8B

Microsoft

B · open · Phi

Phi-4

Microsoft

3:9

overall capability wins (1 ties across 13 scored). Phi-4 leads overall.

  • Reasoning2:4
  • Coding0:3
  • Arena0:1
  • Pricing0:3
  • Speed0:2
  • Specs1:1

Verdict

Phi-4 leads coding (widest gap: +17.0 on SciCode); Phi-4 ranks higher on LMArena (+128 Elo); Phi-3 Mini 3.8B offers ~7.8x larger context (128K vs 16K).

  • Phi-4 leads coding (widest gap: +17.0 on SciCode)
  • Phi-4 ranks higher on LMArena (+128 Elo)
  • Phi-3 Mini 3.8B offers ~7.8x larger context (128K vs 16K)
  • Phi-4 is ~2.6x cheaper on a blended token basis than Phi-3 Mini 3.8B

Cost for 1M in + 250K out

Illustrative chat workload at primary-provider list prices.

Phi-3 Mini 3.8B
$0.26
Phi-4
$0.11

Phi-4 is about 2.5x cheaper on this workload.

Price per 1M tokens

Full comparison

Organization

Microsoft

Microsoft

Family

Phi

Phi

License

MIT

MIT

Open weights

Yes

Yes

Release

Apr 23, 2024

Dec 12, 2024

Knowledge cutoff

-

-

API / provider

Azure AI Foundry

Azure / self-host

Modalities

text → text

text → text

Specs

Context window

128K

16K

A +112K

Max output

4K

4K

tie

Parameters

3.8B

14B

B +10

Pricing

Input $/1M

$0.13

$0.07

B +$0.06

Output $/1M

$0.52

$0.14

B +$0.38

Blended $/1M (3∶1)

$0.23

$0.09

B +$0.14

Speed

tok/s

140

150

B +10

TTFT (s)

0.18

0.15

B +0.03

Reasoning

MMLU-Pro

43.5

70.4

B +26.9 pts

GPQA Diamond

31.9

57.5

B +25.6 pts

Humanity's Last Exam

5.0

3.8

A +1.2 pts

AIME 2025

0.3

18.0

B +17.7 pts

MATH-500

45.7

81.0

B +35.3 pts

Humanity's Last Exam (with tools)

—

—

—

AA-Omniscience Accuracy

—

14.1

—

AA-LCR v1.1

—

—

—

CritPt

—

—

—

MMMU-Pro

—

—

—

IFBench

23.9

23.5

A +0.4 pts

Chartography

—

—

—

Chartography (With Tools)

—

—

—

Coding

SWE-bench Verified

—

—

—

SWE-bench Pro

—

—

—

SWE-bench Multilingual

—

—

—

LiveCodeBench

11.6

23.1

B +11.5 pts

Terminal-Bench 2.1

—

—

—

Aider Polyglot

—

—

—

Terminal-Bench 3

—

—

—

BigCodeBench

29.6

45.5

B +15.9 pts

SciCode

9.0

26.0

B +17.0 pts

CursorBench

—

—

—

SWE-Rebench

—

—

—

NL2Repo-Bench

—

—

—

DeepSWE

—

—

—

WebDev Arena

—

—

—

Terminal-Bench 4.0

—

—

—

CursorBench 4.0

—

—

—

FrontierCode v1.1 (Main)

—

—

—

Arena

LMArena Elo

1,128

1,256

B +128 Elo

Phi-3 Mini 3.8B

Context
128K / 4K out
Parameters
3.8B
Price
$0.13 / $0.52
Speed
140 tok/s · 0.18s TTFT
Modalities
text → text
License
MIT

Original Phi-3 Mini instruct — compact Azure / edge SLM that popularized Microsoft's open Phi line.

Phi-4

Context
16K / 4K out
Parameters
14B
Price
$0.07 / $0.14
Speed
150 tok/s · 0.15s TTFT
Modalities
text → text
License
MIT

Small high-quality dense model from Microsoft - punches above its size on STEM.

Benchmark charts

Winner bars are emphasized. Per-benchmark deltas sit above each chart.

Reasoning & knowledge

MMLU-Pro: B +26.9GPQA: B +25.6HLE: A +1.2AIME: B +17.7MATH-500: B +35.3IFBench: A +0.4

Coding

LiveCode: B +11.5SciCode: B +17.0BigCodeBench: B +15.9

Arena

Arena Elo: B +128

Pick a different pair · Back to catalog