Skip to content
Updated 2h ago

AI Leaderboard - 916 models ranked by capability, price & context

Independent benchmark results from Epoch AI and Artificial Analysis, live API prices and context windows from OpenRouter. Each ranking uses one named measure - we don't blend them into a made-up score.

Highest capability
GPT-6 Astra
ECI 166.6Epoch AI
Leads on reasoning
GPT-6 Astra
95.8% GPQAEpoch AI
Wins at coding
GPT-6 Astra
74.1% DeepSWEEpoch AI
Cheapest in the top 10
GPT-5.6 Sol
$4.00/M blendedOpenRouter
Longest context
Grok 4.20
2.0M tokensOpenRouter
Best open weights
Kimi K3
ECI 157.68Epoch AI

Ranked by the Epoch Capabilities Index (ECI) - Epoch AI's statistical model over dozens of benchmarks. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
201Ministral 3BMistral AI118.125.3———
202Yi-34B01.AI117.314.7———
203phi-3-mini 3.8BMicrosoft117.2————
204Gemini 1.0 ProGoogle117.034.0———
205Stable Beluga 2Stability AI117.0————
206Llama 3.1-8BMeta116.527.0———
207Llama 3-8BMeta116.326.1———
208Qwen2.5-Coder-14B116.3————
209Gemma 3 4BGoogle116.023.2—131K$0.063
210GPT-3.5 Turbo (Jan 2024)OpenAI115.627.2———
211PaLM 2-L114.8————
212Llama 2-70BMeta113.626.3———
213GPT-3.5 Turbo (Jun 2023)OpenAI113.1————
214Qwen2.5-Coder (7B)Alibaba (Qwen)112.9————
215Qwen-14BAlibaba (Qwen)112.8————
216Mistral 7B v0.1Mistral AI112.0————
217Falcon-180BTechnology Innovation Institute111.9————
218internlm-20b111.8————
219Gemma 7BGoogle111.7————
220DeepSeek LLM 67BDeepSeek110.524.6———
221LLaMA-65BMeta109.9————
222Falcon 2 11BTechnology Innovation Institute109.3————
223Mistral 7B v0.3Mistral AI108.715.2———
224DeepSeek-Coder-V2-Lite-Base108.6————
225PaLM 2-M108.0————
226Phi-2Microsoft107.6————
227Qwen2.5-Coder-3B107.4————
228Nemotron-4 15BNVIDIA107.4————
229Yi-9B107.3————
230LLaMA-33BMeta107.1————
231Qwen-7BAlibaba (Qwen)106.5————
232Llama 2-13BMeta105.8————
233PaLM 2-S105.8————
234Llama 2-34BMeta104.8————
235StarCoder 2 15BHugging Face104.6————
236Yi 6B01.AI104.4————
237Falcon-40BTechnology Innovation Institute104.0————
238Baichuan2-13BBaichuan102.8————
239Qwen2.5-Coder (1.5B)Alibaba (Qwen)102.6————
240internlm-7b102.4————
241Llama 3.2 1BMeta102.023.9———
242INTELLECT-1Prime Intellect,Hugging Face,Arcee AI100.4————
243MPT-30BMosaicML100.2————
244LLaMA-13BMeta100.1————
245Llama 2-7BMeta98.5————
246chatglm2-6b98.4————
247LLaMA-7BMeta96.1————
248Baichuan 2-7BBaichuan95.8————
249DeepSeek Coder 33BDeepSeek95.7————
250Falcon-7BTechnology Innovation Institute94.5————
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research