Skip to content

Best for reasoning

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up.

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
1Kimi K3Moonshot AI · +1 variant157.793.168.51.05M$6.00
2GLM 5.2Z.ai (Zhipu AI)151.891.943.81.05M$1.39
3DeepSeek V4 Pro 0813DeepSeek155.491.7—1.05M$1.41
4DeepSeek V4 Flash 0731DeepSeek154.591.0—1.31M$0.093
5GLM 5.3Z.ai (Zhipu AI)155.690.969.01.31M$2.15
6MiniMax M3MiniMax147.090.9—1.05M$0.53
7DeepSeek v4 ProhighDeepSeek · +1 variant—90.9———
8Kimi K2.6Moonshot AI151.090.8—262K$1.34
9GLM 5.3 FlashZ.ai (Zhipu AI)151.990.263.41.31M$0.24
10GLM 5.1Z.ai (Zhipu AI)149.989.9—205K$2.15
11Inkling SmallxhighThinking Machines Lab—88.5———
12InklingxhighThinking Machines Lab—88.3———
13Kimi K2.7 CodeMoonshot AI150.087.930.5262K$1.32
14GLM 5Z.ai (Zhipu AI)145.887.8—205K$0.93
15Kimi K2.5 (Fireworks)Moonshot AI—87.6———
16Qwen3.5 397B A17BAlibaba (Qwen)146.686.4—262K$1.29
17Qwen3.6 27BAlibaba (Qwen)146.585.9—262K$1.04
18Nemotron 3 UltraNVIDIA146.285.4—262K$1.05
19Qwen3.6 35B A3BAlibaba (Qwen)—84.8—262K$0.36
20Kimi K2 Thinking TurboMoonshot AI—84.2———
21Qwen3.5-35B-A3BAlibaba (Qwen)142.583.5—262K$0.45
22DeepSeek V3.2DeepSeek146.383.4—164K$0.32
23GLM 4.7Z.ai (Zhipu AI)143.583.3—205K$1.00
24Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba (Qwen)143.980.1———
25Qwen3.5-9BAlibaba (Qwen)139.479.0—262K$0.11
26DeepSeek-R1 (May 2025)DeepSeek141.376.3———
27Gemma 4 31BGoogle142.775.8—262K$0.15
28gpt-oss-120bOpenAI139.975.8—131K$0.070
29Gemma 4 26B A4BGoogle141.873.2—262K$0.14
30R1DeepSeek139.071.7—64K$1.15
31Qwen3 235B A22BAlibaba (Qwen)139.470.7—131K$0.80
32Qwen3-30B-A3B-Thinking (Jul 2025)Alibaba (Qwen)139.670.1———
33DeepSeek-V3 (Mar 2025)DeepSeek135.967.6———
34Llama 4 Maverick (FP8)Meta—67.0———
35Qwen3 32BAlibaba (Qwen)138.565.7—131K$0.13
36QwQ-32BAlibaba (Qwen)137.665.3———
37DeepSeek-R1-Distill-Qwen-32BDeepSeek137.464.1———
38Qwen3 14BAlibaba (Qwen)138.263.8—131K$0.15
39Qwen3 30B A3BAlibaba (Qwen)136.261.7—131K$0.23
40gpt-oss-20bOpenAI137.860.8—131K$0.036
41GLM 4.7 FlashZ.ai (Zhipu AI)—60.5—200K$0.15
42Qwen3 8BAlibaba (Qwen)136.256.8—131K$0.20
43DeepSeek V3DeepSeek132.456.5—164K$0.45
44Magistral Small 1.0Mistral AI133.256.1———
45Phi 4Microsoft130.456.1—16K$0.087
46DeepSeek-R1-Distill-Llama-70BDeepSeek—55.7———
47Qwen3-30B-A3B-Instruct (Jul 2025)Alibaba (Qwen)137.455.6———
48Qwen3-4BAlibaba (Qwen)—52.3———
49Llama 4 ScoutMeta129.651.8—1.31M$0.15
50Mistral Large 2 (Nov 2024)Mistral AI128.551.3———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research