Skip to content

Open-weight leaderboard

Models whose weights are published, ranked by the Epoch Capabilities Index.

Models whose weights are published, ranked by the Epoch Capabilities Index. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
1Kimi K3Moonshot AI157.793.168.51.05M$6.00
2GLM 5.3Z.ai (Zhipu AI)155.690.969.01.31M$2.15
3DeepSeek V4 Pro 0813DeepSeek155.491.7—1.05M$1.41
4DeepSeek V4.1 FlashNEWDeepSeek155.0——1.05M$0.53
5DeepSeek V4 Flash 0731DeepSeek154.591.0—1.31M$0.093
6GLM 5.3 FlashZ.ai (Zhipu AI)151.990.263.41.31M$0.24
7GLM 5.2Z.ai (Zhipu AI)151.891.943.81.05M$1.51
8Kimi K2.6Moonshot AI151.090.8—262K$1.34
9Inkling SmallThinking Machines Lab150.1——524K$0.64
10Kimi K2.7 CodeMoonshot AI150.087.930.5262K$1.32
11GLM 5.1Z.ai (Zhipu AI)149.989.9—205K$2.15
12Qwen 3.8 27BAlibaba (Qwen)149.4————
13DeepSeek V4 Pro 0423DeepSeek149.1——1.05M$1.19
14InklingThinking Machines Lab148.6——524K$1.76
15Kimi K2.5Moonshot AI148.1——262K$0.90
16MiniMax M3MiniMax147.090.9—1.05M$0.53
17MiniMax M2.5MiniMax146.7——205K$0.47
18Qwen3.5 397B A17BAlibaba (Qwen)146.686.4—262K$1.29
19Qwen3.6 27BAlibaba (Qwen)146.585.9—262K$1.04
20DeepSeek V3.2DeepSeek146.383.4—164K$0.32
21Nemotron 3 UltraNVIDIA146.285.4—262K$1.05
22DeepSeek V4 Flash 0423DeepSeek146.1——1.05M$0.17
23Kimi K2 ThinkingMoonshot AI146.0——262K$1.07
24GLM 5Z.ai (Zhipu AI)145.887.8—205K$0.93
25MiniMax M2.7MiniMax145.8——205K$0.37
26DeepSeek V3.2 ExpDeepSeek145.0——164K$0.30
27Qwen 3.6 35B-A3BAlibaba (Qwen)143.9————
28Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba (Qwen)143.980.1———
29GLM 4.7Z.ai (Zhipu AI)143.583.3—205K$1.00
30Gemma 4 31BGoogle142.775.8—262K$0.15
31Qwen3.5-35B-A3BAlibaba (Qwen)142.583.5—262K$0.45
32Gemma 4 26B A4BGoogle141.873.2—262K$0.14
33Mistral Medium 3.5Mistral AI141.4——262K$3.00
34DeepSeek-R1 (May 2025)DeepSeek141.376.3———
35Kimi K2 (Jul 2025)Moonshot AI140.1————
36gpt-oss-120bOpenAI139.975.8—131K$0.070
37DeepSeek V3.1DeepSeek139.9——164K$0.42
38Qwen3-30B-A3B-Thinking (Jul 2025)Alibaba (Qwen)139.670.1———
39Qwen3.5-9BAlibaba (Qwen)139.479.0—262K$0.11
40Qwen3 235B A22BAlibaba (Qwen)139.470.7—131K$0.80
41R1DeepSeek139.071.7—64K$1.15
42Qwen3-235B-A22B-Instruct (Jul 2025)Alibaba (Qwen)138.9————
43Qwen3 32BAlibaba (Qwen)138.565.7—131K$0.13
44Qwen3 14BAlibaba (Qwen)138.263.8—131K$0.15
45gpt-oss-20bOpenAI137.860.8—131K$0.036
46QwQ-32BAlibaba (Qwen)137.665.3———
47DeepSeek-R1-Distill-Qwen-32BDeepSeek137.464.1———
48Qwen3-30B-A3B-Instruct (Jul 2025)Alibaba (Qwen)137.455.6———
49Qwen3 30B A3BAlibaba (Qwen)136.261.7—131K$0.23
50Qwen3 8BAlibaba (Qwen)136.256.8—131K$0.20
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research