Skip to content

Best for math

Ranked by Epoch's mock AIME 2024-2025 (competition mathematics).

Ranked by Epoch's mock AIME 2024-2025 (competition mathematics). Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
1DeepSeek V4 Pro 0813DeepSeek155.491.7—1.05M$1.41
2Kimi K3Moonshot AI · +1 variant157.793.168.51.05M$6.00
3DeepSeek v4 ProhighDeepSeek · +1 variant—90.9———
4Kimi K2.6Moonshot AI151.090.8—262K$1.34
5Kimi K2.7 CodeMoonshot AI150.087.930.5262K$1.32
6DeepSeek V4 Flash 0731DeepSeek154.591.0—1.31M$0.093
7GLM 5.3 FlashZ.ai (Zhipu AI)151.990.263.41.31M$0.24
8GLM 5.1Z.ai (Zhipu AI)149.989.9—205K$2.15
9Kimi K2.5 (Fireworks)Moonshot AI—87.6———
10GLM 5.3Z.ai (Zhipu AI)155.690.969.01.31M$2.15
11Qwen3.6 27BAlibaba (Qwen)146.585.9—262K$1.04
12Inkling SmallxhighThinking Machines Lab—88.5———
13Qwen3.5 397B A17BAlibaba (Qwen)146.686.4—262K$1.29
14gpt-oss-120bOpenAI139.975.8—131K$0.070
15InklingxhighThinking Machines Lab—88.3———
16DeepSeek V3.2DeepSeek146.383.4—164K$0.32
17Nemotron 3 UltraNVIDIA146.285.4—262K$1.05
18Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba (Qwen)143.980.1———
19Qwen3.6 35B A3BAlibaba (Qwen)—84.8—262K$0.36
20GLM 5.2Z.ai (Zhipu AI)151.891.943.81.05M$1.51
21GLM 4.7Z.ai (Zhipu AI)143.583.3—205K$1.00
22Kimi K2 Thinking TurboMoonshot AI—84.2———
23Gemma 4 26B A4BGoogle141.873.2—262K$0.14
24GLM 5Z.ai (Zhipu AI)145.887.8—205K$0.93
25Gemma 4 31BGoogle142.775.8—262K$0.15
26MiniMax M3MiniMax147.090.9—1.05M$0.53
27Qwen3-30B-A3B-Thinking (Jul 2025)Alibaba (Qwen)139.670.1———
28Qwen3.5-35B-A3BAlibaba (Qwen)142.583.5—262K$0.45
29Qwen3 32BAlibaba (Qwen)138.565.7—131K$0.13
30DeepSeek-R1 (May 2025)DeepSeek141.376.3———
31Qwen3 14BAlibaba (Qwen)138.263.8—131K$0.15
32gpt-oss-20bOpenAI137.860.8—131K$0.036
33Qwen3 30B A3BAlibaba (Qwen)136.261.7—131K$0.23
34Qwen3-30B-A3B-Instruct (Jul 2025)Alibaba (Qwen)137.455.6———
35Qwen3.5-9BAlibaba (Qwen)139.479.0—262K$0.11
36QwQ-32BAlibaba (Qwen)137.665.3———
37GLM 4.7 FlashZ.ai (Zhipu AI)—60.5—200K$0.15
38Qwen3 8BAlibaba (Qwen)136.256.8—131K$0.20
39Qwen3.5-4BAlibaba (Qwen)—————
40DeepSeek-R1-Distill-Qwen-32BDeepSeek137.464.1———
41R1DeepSeek139.071.7—64K$1.15
42DeepSeek-R1-Distill-Llama-70BDeepSeek—55.7———
43DeepSeek-R1-Distill-Qwen-14BDeepSeek135.444.7———
44DeepSeek-V3 (Mar 2025)DeepSeek135.967.6———
45Mistral Small 3.2Mistral AI131.749.1———
46Magistral Small 1.0Mistral AI133.256.1———
47Magistral Small 1.2Mistral AI131.447.6———
48Gemma 3 27BGoogle130.047.7—131K$0.17
49DeepSeek-R1-Distill-Qwen-1.5BDeepSeek—33.6———
50Llama 4 Maverick (FP8)Meta—67.0———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research