LLMcompare

Search

Search for a command to run...

Coding

WebDev Arena

LMArena Code / WebDev human-preference Elo for front-end and agentic web development tasks; this is a time-varying leaderboard snapshot, not a fixed model property.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Coding
Metric
Elo
Direction
Higher is better
Coverage
74 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Claude Fable 5.1

    Anthropicclosed

    1,758
  2. 2
    Claude Opus 5

    Anthropicclosed

    1,687
  3. 3
    Qwen3.8 Max 0902

    Alibabaclosed

    1,681
  4. 4
    Kimi K3

    Moonshotopen

    1,674
  5. 5
    Qwen3.8 Max

    Alibabaopen

    1,671
  6. 61,635
  7. 7
    Claude Fable 5

    Anthropicclosed

    1,628
  8. 8
    Hy4 Preview

    Tencentopen

    1,624
  9. 91,620
  10. 10
    Grok 4.6

    xAIclosed

    1,618
  11. 11
    GPT-5.6 Sol

    OpenAIclosed

    1,617
  12. 12
    GLM-5.3

    Zhipu AIopen

    1,614
  13. 13
    GLM-5.3 Flash

    Zhipu AIopen

    1,607
  14. 14
    Qwen3.8 27B

    Alibabaopen

    1,593
  15. 15
    GLM-5.2

    Zhipu AIopen

    1,592
  16. 16
    Gemini 3.7 Flash

    Googleclosed

    1,588
  17. 17
    DeepSeek V4 Flash

    DeepSeekopen

    1,580
  18. 18
    Gemini 3.8 Flash

    Googleclosed

    1,567
  19. 19
    Claude Opus 4.7

    Anthropicclosed

    1,557
  20. 20
    Grok 4.5

    xAIclosed

    1,555
  21. 21
    Muse Spark 1.1

    Metaclosed

    1,542
  22. 22
    Claude Opus 4.8

    Anthropicclosed

    1,539
  23. 23
    Claude Opus 4.6

    Anthropicclosed

    1,537
  24. 24
    Gemini 3.6 Flash

    Googleclosed

    1,537
  25. 25
    Claude Sonnet 5

    Anthropicclosed

    1,537
  26. 26
    Muse Spark 1.2

    Metaclosed

    1,535
  27. 27
    GPT-5.6 Terra

    OpenAIclosed

    1,521
  28. 28
    Claude Sonnet 4.6

    Anthropicclosed

    1,521
  29. 29
    GPT-5.6 Luna

    OpenAIclosed

    1,519
  30. 30
    Qwen3.7 Max

    Alibabaclosed

    1,517
  31. 31
    Hy3

    Tencentopen

    1,513
  32. 32
    GPT-5.5

    OpenAIclosed

    1,510
  33. 33
    Kimi K2.6

    Moonshotclosed

    1,509
  34. 34
    GLM-5.1

    Zhipu AIopen

    1,508
  35. 35
    Gemini 3.5 Flash

    Googleclosed

    1,492
  36. 36
    MiniMax M3

    MiniMaxopen

    1,487
  37. 37
    MiMo-V2.5-Pro

    Xiaomiopen

    1,475
  38. 38
    Kimi K2.7 Code

    Moonshotopen

    1,473
  39. 39
    Claude Opus 4.5

    Anthropicclosed

    1,468
  40. 40
    Qwen3.6 Plus

    Alibabaclosed

    1,461
  41. 41
    Gemini 3.1 Pro

    Googleclosed

    1,447
  42. 421,447
  43. 43
    DeepSeek V4 Pro

    DeepSeekopen

    1,446
  44. 44
    Gemini 3 Pro

    Googleclosed

    1,439
  45. 45
    Gemini 3 Flash

    Googleclosed

    1,438
  46. 46
    MiMo-V2.5

    Xiaomiopen

    1,437
  47. 47
    GLM-5

    Zhipu AIopen

    1,436
  48. 48
    GLM-4.7

    Zhipu AIopen

    1,435
  49. 49
    Inkling

    Thinking Machinesopen

    1,410
  50. 50
    Inkling-Small

    Thinking Machinesopen

    1,407
  51. 511,399
  52. 52
    MiniMax M2.7

    MiniMaxclosed

    1,398
  53. 53
    GPT-5.4 Mini

    OpenAIclosed

    1,397
  54. 54
    Claude Opus 4.1

    Anthropicclosed

    1,389
  55. 55
    GPT-5.4

    OpenAIclosed

    1,387
  56. 56
    Claude Sonnet 4.5

    Anthropicclosed

    1,386
  57. 57
    MiniMax M2.5

    MiniMaxopen

    1,384
  58. 58
    Solar Pro 4

    Upstageclosed

    1,372
  59. 59
    Gemma 4 31B

    Googleopen

    1,364
  60. 60
    Gemma 4 26B

    Googleopen

    1,362
  61. 61
    Muse Glimmer

    Metaopen

    1,360
  62. 62
    Grok 4.3

    xAIclosed

    1,357
  63. 63
    Hunyuan 3 Preview

    Tencentclosed

    1,356
  64. 64
    Claude Haiku 4.5

    Anthropicclosed

    1,329
  65. 65
    DeepSeek V3.2

    DeepSeekopen

    1,325
  66. 661,265
  67. 671,254
  68. 68
    Qwen3.5 Flash

    Alibabaclosed

    1,238
  69. 691,237
  70. 70
    Mistral Large 3

    Mistralopen

    1,230
  71. 71
    Gemini 2.5 Pro

    Googleclosed

    1,226
  72. 72
    Devstral 2

    Mistralopen

    1,194
  73. 731,192
  74. 74
    Mercury 2

    Inception Labsclosed

    1,167

Back to all benchmarks