Skip to content

Open-weight leaderboard

Models whose weights are published, ranked by the Epoch Capabilities Index.

Models whose weights are published, ranked by the Epoch Capabilities Index. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
51DeepSeek-V3 (Mar 2025)DeepSeek135.967.6———
52DeepSeek-R1-Distill-Qwen-14BDeepSeek135.444.7———
53Magistral Small 1.0Mistral AI133.256.1———
54DeepSeek V3DeepSeek132.456.5—164K$0.45
55Llama 4 MaverickMeta132.2——1.05M$0.30
56Mistral Small 3.2Mistral AI131.749.1———
57Magistral Small 1.2Mistral AI131.447.6———
58Phi 4Microsoft130.456.1—16K$0.087
59Gemma 3 27BGoogle130.047.7—131K$0.17
60Llama 4 ScoutMeta129.651.8—1.31M$0.15
61Qwen2.5-72BAlibaba (Qwen)129.049.1———
62Llama 3.1-405BMeta128.850.9———
63Mistral Large 2 (Nov 2024)Mistral AI128.551.3———
64Qwen2.5-32BAlibaba (Qwen)128.546.1———
65Mistral Large 2 (Jul 2024)Mistral AI127.549.0———
66Mistral Small 3.1Mistral AI127.547.5———
67Llama 3.3 70BMeta127.347.4———
68Mistral Small 3Mistral AI127.147.3—33K$0.058
69Llama 3.1-70BMeta125.944.2———
70Llama 3.2 90BMeta125.541.0———
71Qwen2-72BAlibaba (Qwen)125.340.8———
72DeepSeek-V2 (MoE-236B, May 2024)DeepSeek124.8————
73Gemma 3 12BGoogle123.539.5—131K$0.075
74Llama 3-70BMeta122.940.6———
75Gemma 2 27BGoogle122.136.5—8K$0.65
76Mixtral 8x22BMistral AI122.034.1———
77phi-3-small 7.4BMicrosoft121.8————
78phi-3-medium 14BMicrosoft121.227.6———
79Gemma 2 9BGoogle119.827.5———
80Qwen2.5-Coder-32BAlibaba (Qwen)119.4————
81Command R+Cohere119.3————
82Mistral NemoMistral AI118.629.9—131K$0.022
83Qwen2.5-7BAlibaba (Qwen)118.435.5———
84Mixtral 8x7BMistral AI118.430.6———
85Yi-34B01.AI117.314.7———
86phi-3-mini 3.8BMicrosoft117.2————
87Stable Beluga 2Stability AI117.0————
88Llama 3.1-8BMeta116.527.0———
89Llama 3-8BMeta116.326.1———
90Gemma 3 4BGoogle116.023.2—131K$0.063
91Llama 2-70BMeta113.626.3———
92Qwen2.5-Coder (7B)Alibaba (Qwen)112.9————
93Qwen-14BAlibaba (Qwen)112.8————
94Mistral 7B v0.1Mistral AI112.0————
95Falcon-180BTechnology Innovation Institute111.9————
96Gemma 7BGoogle111.7————
97DeepSeek LLM 67BDeepSeek110.524.6———
98LLaMA-65BMeta109.9————
99Falcon 2 11BTechnology Innovation Institute109.3————
100Mistral 7B v0.3Mistral AI108.715.2———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research