Skip to content
Updated 2h ago

AI Leaderboard - 916 models ranked by capability, price & context

Independent benchmark results from Epoch AI and Artificial Analysis, live API prices and context windows from OpenRouter. Each ranking uses one named measure - we don't blend them into a made-up score.

Highest capability
GPT-6 Astra
ECI 166.6Epoch AI
Leads on reasoning
GPT-6 Astra
95.8% GPQAEpoch AI
Wins at coding
GPT-6 Astra
74.1% DeepSWEEpoch AI
Cheapest in the top 10
GPT-5.6 Sol
$4.00/M blendedOpenRouter
Longest context
Grok 4.20
2.0M tokensOpenRouter
Best open weights
Kimi K3
ECI 157.68Epoch AI

Ranked by the Epoch Capabilities Index (ECI) - Epoch AI's statistical model over dozens of benchmarks. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
151Phi 4Microsoft130.456.1—16K$0.087
152Gemma 3 27BGoogle130.047.7—131K$0.17
153Claude 3.5 SonnetAnthropic130.0————
154Llama 4 ScoutMeta129.651.8—1.31M$0.15
155GPT-4.1 NanoOpenAI129.648.9—1.05M$0.17
156Gemini 1.5 Flash (Sep 2024)Google129.447.3———
157Qwen2.5-72BAlibaba (Qwen)129.049.1———
158GPT-4o (May 2024)OpenAI129.048.9———
159GPT-4o (Nov 2024)OpenAI128.847.9———
160GPT-4o (Aug 2024)OpenAI128.849.2———
161Llama 3.1-405BMeta128.850.9———
162Mistral Large 2 (Nov 2024)Mistral AI128.551.3———
163Qwen2.5-32BAlibaba (Qwen)128.546.1———
164Mistral Large 2 (Jul 2024)Mistral AI127.549.0———
165Mistral Small 3.1Mistral AI127.547.5———
166Llama 3.3 70BMeta127.347.4———
167GPT-4 Turbo (Apr 2024)OpenAI127.346.6———
168Claude 3.5 HaikuAnthropic127.2————
169Mistral Small 3Mistral AI127.147.3—33K$0.058
170Claude 3 OpusAnthropic126.947.2———
171Gemini 1.5 Pro (May 2024)Google126.945.9———
172GPT-4o-miniOpenAI126.637.7—128K$0.26
173GPT-4 Turbo (Nov 2023)OpenAI126.5————
174Llama 3.1-70BMeta125.944.2———
175GPT-4 (Mar 2023)OpenAI125.935.7———
176Llama 3.2 90BMeta125.541.0———
177Qwen2-72BAlibaba (Qwen)125.340.8———
178DeepSeek-V2 (MoE-236B, May 2024)DeepSeek124.8————
179Amazon Nova ProAmazon123.8————
180Gemma 3 12BGoogle123.539.5—131K$0.075
181GPT-4 (Jun 2023)OpenAI123.130.7———
182Llama 3-70BMeta122.940.6———
183Gemini 1.5 Flash (May 2024)Google122.640.4———
184Gemma 2 27BGoogle122.136.5—8K$0.65
185Mistral LargeMistral AI122.038.8—128K$3.00
186Mixtral 8x22BMistral AI122.034.1———
187phi-3-small 7.4BMicrosoft121.8————
188phi-3-medium 14BMicrosoft121.227.6———
189Claude 3 SonnetAnthropic120.740.6———
190Claude InstantAnthropic120.2————
191Claude 2Anthropic120.134.7———
192Gemma 2 9BGoogle119.827.5———
193Qwen2.5-Coder-32BAlibaba (Qwen)119.4————
194Command R+Cohere119.3————
195Claude 2.1Anthropic119.233.0———
196Mistral NemoMistral AI118.629.9—131K$0.022
197GPT-3.5 Turbo (Nov 2023)OpenAI118.528.0———
198Qwen2.5-7BAlibaba (Qwen)118.435.5———
199Mixtral 8x7BMistral AI118.430.6———
200Claude 3 HaikuAnthropic118.336.3———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research