LLMcompare

Search

Search for a command to run...

Reasoning & knowledge

AA-LCR v1.1

Artificial Analysis Long Context Reasoning: 100 open-answer questions requiring multi-step synthesis across documents of 10k-100k tokens (academic papers, financial reports, legal and government documents). Pass/fail graded by an LLM judge against official answers; scores are tied to the AA-LCR version and grader model and are not a needle-in-a-haystack retrieval measure.

How models are tested

Scores are published numbers from the sources below. They are only comparable when the evaluation version, tools, agent harness, and sampling setup match. Missing scores are omitted rather than treated as zero.

Category
Reasoning & knowledge
Metric
Percent
Direction
Higher is better
Coverage
151 / 318Models in this catalog with a published score
Updated
Catalog snapshot date

Scores

  1. 1
    Kimi K3

    Moonshotopen

    88.7%
  2. 2
    GPT-5.5

    OpenAIclosed

    84.3%
  3. 3
    GPT-6 Sol

    OpenAIclosed

    84.0%
  4. 4
    GPT-5.6 Sol

    OpenAIclosed

    84.0%
  5. 584.0%
  6. 6
    GPT-5.6 Luna

    OpenAIclosed

    83.7%
  7. 7
    Muse Glimmer

    Metaopen

    83.3%
  8. 8
    GPT-6 Luna

    OpenAIclosed

    83.0%
  9. 9
    GPT-5.6 Terra

    OpenAIclosed

    83.0%
  10. 10
    Claude Fable 5.1

    Anthropicclosed

    83.0%
  11. 11
    Muse Spark 1.3

    Metaclosed

    83.0%
  12. 12
    MiniMax M3

    MiniMaxopen

    83.0%
  13. 13
    Claude Fable 5

    Anthropicclosed

    82.3%
  14. 14
    GPT-5.4

    OpenAIclosed

    82.0%
  15. 15
    Claude Sonnet 5

    Anthropicclosed

    82.0%
  16. 16
    Qwen3.8 27B

    Alibabaopen

    82.0%
  17. 17
    Gemini 3.7 Flash

    Googleclosed

    81.7%
  18. 18
    Gemini 3.8 Flash

    Googleclosed

    81.3%
  19. 19
    GPT-6 Astra

    OpenAIclosed

    81.0%
  20. 20
    Kimi K2.6

    Moonshotclosed

    81.0%
  21. 21
    Grok 4.6

    xAIclosed

    80.3%
  22. 22
    DeepSeek V4 Pro 0813

    DeepSeekclosed

    80.3%
  23. 23
    Qwen3.8 Max

    Alibabaopen

    80.3%
  24. 24
    Claude Sonnet 4.6

    Anthropicclosed

    80.0%
  25. 25
    Gemini 3.6 Flash

    Googleclosed

    80.0%
  26. 26
    K2 Horizon 375B A23B

    MBZUAI Institute of Foundation Modelsopen

    80.0%
  27. 27
    GLM-5.3 Flash

    Zhipu AIopen

    80.0%
  28. 2879.7%
  29. 2979.7%
  30. 30
    MiMo-V2.5-Pro

    Xiaomiopen

    79.7%
  31. 31
    GLM-5.3

    Zhipu AIopen

    79.7%
  32. 32
    Claude Opus 5

    Anthropicclosed

    79.3%
  33. 33
    Grok 4.5

    xAIclosed

    79.3%
  34. 34
    Kimi K2.7 Code

    Moonshotopen

    79.3%
  35. 3579.3%
  36. 36
    Muse Spark 1.2

    Metaclosed

    79.0%
  37. 37
    Qwen3.7 Max

    Alibabaclosed

    79.0%
  38. 38
    Hy3

    Tencentopen

    79.0%
  39. 39
    Claude Opus 4.7

    Anthropicclosed

    78.7%
  40. 40
    Qwen3.6 Plus

    Alibabaclosed

    78.3%
  41. 41
    GLM-5.2

    Zhipu AIopen

    78.3%
  42. 42
    MiniMax M2.7

    MiniMaxclosed

    78.3%
  43. 43
    GPT-5

    OpenAIclosed

    78.2%
  44. 44
    Gemini 3 Flash

    Googleclosed

    78.0%
  45. 45
    Kimi K2.5

    Moonshotopen

    78.0%
  46. 46
    Claude Opus 4.6

    Anthropicclosed

    78.0%
  47. 47
    Claude Opus 4.8

    Anthropicclosed

    77.7%
  48. 48
    Muse Spark 1.1

    Metaclosed

    77.7%
  49. 49
    Claude Opus 4.5

    Anthropicclosed

    77.3%
  50. 50
    Qwen3.6 27B

    Alibabaopen

    77.3%
  51. 5177.3%
  52. 52
    Inkling

    Thinking Machinesopen

    77.3%
  53. 53
    Grok 4.7

    xAIclosed

    77.0%
  54. 54
    GPT-5.4 Mini

    OpenAIclosed

    77.0%
  55. 55
    GPT-5.4 Nano

    OpenAIclosed

    76.7%
  56. 5676.0%
  57. 57
    GLM-5

    Zhipu AIopen

    75.7%
  58. 58
    Inkling-Small

    Thinking Machinesopen

    75.7%
  59. 59
    o3

    OpenAIclosed

    74.7%
  60. 60
    DeepSeek V4 Pro

    DeepSeekopen

    74.7%
  61. 61
    DeepSeek V4 Flash

    DeepSeekopen

    74.3%
  62. 6274.3%
  63. 63
    Solar Pro 4

    Upstageclosed

    74.0%
  64. 64
    GLM-5.1

    Zhipu AIopen

    73.7%
  65. 65
    Grok 4 Fast

    xAIclosed

    73.7%
  66. 66
    DeepSeek V3.2

    DeepSeekopen

    73.3%
  67. 67
    MiniMax M2.5

    MiniMaxopen

    73.3%
  68. 68
    Gemini 3.5 Flash

    Googleclosed

    73.3%
  69. 69
    Grok 4.3

    xAIclosed

    73.0%
  70. 70
    Qwen3.7 Plus

    Alibabaclosed

    73.0%
  71. 71
    MiMo-V2.5

    Xiaomiopen

    73.0%
  72. 72
    Ling-3.0 Flash

    InclusionAIopen

    73.0%
  73. 73
    GPT-5 Mini

    OpenAIclosed

    72.3%
  74. 74
    Qwen3.6 35B A3B

    Alibabaopen

    71.7%
  75. 75
    MiMo-V2-Flash

    Xiaomiopen

    71.3%
  76. 76
    GLM-4.7

    Zhipu AIopen

    71.0%
  77. 77
    Ring-2.6-1T

    InclusionAIopen

    70.0%
  78. 78
    Gemma 4 31B

    Googleopen

    69.7%
  79. 7969.3%
  80. 80
    Gemini 2.5 Pro

    Googleclosed

    69.0%
  81. 8168.3%
  82. 82
    GPT-4.1

    OpenAIclosed

    68.3%
  83. 83
    Grok 4

    xAIclosed

    68.0%
  84. 84
    Gemini 2.5 Flash

    Googleclosed

    65.3%
  85. 85
    LongCat-2.0

    Meituanopen

    65.0%
  86. 86
    o1

    OpenAIclosed

    65.0%
  87. 87
    Hunyuan 3 Preview

    Tencentclosed

    64.7%
  88. 8863.7%
  89. 89
    Seed-OSS-36B-Instruct

    ByteDance Seedopen

    61.3%
  90. 90
    o4-mini

    OpenAIclosed

    61.0%
  91. 9160.3%
  92. 92
    Ling-3.0 Tiny

    InclusionAIopen

    60.3%
  93. 93
    Grok 3

    xAIclosed

    58.0%
  94. 94
    DeepSeek R1

    DeepSeekopen

    57.7%
  95. 95
    K2 Think V2

    MBZUAI / G42 / LLM360open

    57.0%
  96. 96
    DeepSeek V3.1

    DeepSeekopen

    56.7%
  97. 97
    DeepSeek R1-0528

    DeepSeekopen

    55.7%
  98. 9855.7%
  99. 99
    Kimi K2

    Moonshotopen

    53.0%
  100. 100
    gpt-oss-120B

    OpenAIopen

    52.0%
  101. 10150.0%
  102. 102
    Qwen3 Max

    Alibabaclosed

    50.0%
  103. 103
    Step-3.5 Flash

    StepFunclosed

    50.0%
  104. 104
    Mistral Small 4

    Mistralopen

    49.7%
  105. 10549.0%
  106. 106
    Qwen3 Coder 480B

    Alibabaopen

    45.7%
  107. 107
    GPT-5 Nano

    OpenAIclosed

    45.0%
  108. 10845.0%
  109. 109
    GPT-4.1 Mini

    OpenAIclosed

    44.0%
  110. 110
    Mercury 2

    Inception Labsclosed

    43.7%
  111. 111
    Mistral Medium 3.1

    Mistralclosed

    42.7%
  112. 112
    GLM-4.7 Flash

    Zhipu AIclosed

    41.7%
  113. 113
    Ling-2.6 1T

    InclusionAIopen

    41.7%
  114. 114
    INTELLECT-3

    Prime Intellectopen

    39.7%
  115. 11538.0%
  116. 116
    Mistral Large 3

    Mistralopen

    36.0%
  117. 117
    gpt-oss-20B

    OpenAIopen

    34.7%
  118. 11832.7%
  119. 119
    Devstral 2

    Mistralopen

    32.3%
  120. 120
    Solar Pro 3

    Upstageopen

    32.3%
  121. 121
    Gemini 2.0 Flash

    Googleclosed

    31.3%
  122. 122
    Ling-2.6 Flash

    InclusionAIclosed

    31.3%
  123. 123
    DeepSeek V3

    DeepSeekopen

    29.3%
  124. 12428.0%
  125. 12527.7%
  126. 126
    Claude 3 Haiku

    Anthropicclosed

    27.7%
  127. 127
    QwQ-32B

    Alibabaopen

    26.7%
  128. 128
    Ministral 3 14B

    Mistralopen

    26.3%
  129. 129
    Ministral 3 8B

    Mistralopen

    25.7%
  130. 13025.3%
  131. 13124.3%
  132. 132
    Command A

    Cohereclosed

    21.3%
  133. 133
    GPT-4.1 Nano

    OpenAIclosed

    20.3%
  134. 13420.3%
  135. 135
    Devstral Small

    Mistralopen

    18.7%
  136. 136
    Llama 3.1 8B

    Metaopen

    18.0%
  137. 137
    Ministral 3 3B

    Mistralopen

    17.0%
  138. 13815.7%
  139. 139
    Phi-4-mini

    Microsoftopen

    15.3%
  140. 14013.3%
  141. 14112.7%
  142. 14210.3%
  143. 14310.0%
  144. 1448.7%
  145. 145
    Gemma 3 12B

    Googleopen

    8.3%
  146. 1468.3%
  147. 147
    Gemma 3 27B

    Googleopen

    7.3%
  148. 1487.0%
  149. 149
    Gemma 3 4B

    Googleopen

    6.7%
  150. 150
    Llama 3.2 1B

    Metaopen

    6.7%
  151. 151
    Llama 3.2 3B

    Metaopen

    4.3%

Back to all benchmarks