Skip to content
Updated 30m ago

AI Leaderboard - 916 models ranked by capability, price & context

Independent benchmark results from Epoch AI and Artificial Analysis, live API prices and context windows from OpenRouter. Each ranking uses one named measure - we don't blend them into a made-up score.

Highest capability
GPT-6 Astra
ECI 166.6Epoch AI
Leads on reasoning
GPT-6 Astra
95.8% GPQAEpoch AI
Wins at coding
GPT-6 Astra
74.1% DeepSWEEpoch AI
Cheapest in the top 10
GPT-5.6 Sol
$4.00/M blendedOpenRouter
Longest context
Grok 4.20
2.0M tokensOpenRouter
Best open weights
Kimi K3
ECI 157.68Epoch AI

Ranked by the Epoch Capabilities Index (ECI) - Epoch AI's statistical model over dozens of benchmarks. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
51GPT-5.1OpenAI149.787.6—400K$3.44
52Qwen 3.8 27BAlibaba (Qwen)149.4————
53Qwen 3.6 Max (Preview)Alibaba (Qwen)149.387.4———
54DeepSeek V4 Pro 0423DeepSeek149.1——1.05M$1.19
55Grok 4.3 BetaxAI149.188.8———
56GPT-5.4 MiniOpenAI148.883.6—400K$1.69
57InklingThinking Machines Lab148.6——524K$1.76
58Kimi K2.5Moonshot AI148.1——262K$0.90
59Qwen 3.6 PlusAlibaba (Qwen)147.7————
60o3 ProOpenAI147.4——200K$35.00
61Qwen3.7 PlusAlibaba (Qwen)147.487.9—1M$0.56
62MiniMax M3MiniMax147.090.9—1.05M$0.53
63o3OpenAI146.981.8—200K$3.50
64Claude Sonnet 4.5Anthropic146.8——1M$6.00
65Qwen 3.5 Plus (hosted 397B-A17B)Alibaba (Qwen)146.8————
66MiniMax M2.5MiniMax146.7——205K$0.47
67Qwen3.5 397B A17BAlibaba (Qwen)146.686.4—262K$1.29
68Qwen3.6 27BAlibaba (Qwen)146.585.9—262K$1.04
69Grok 4xAI146.587.0———
70DeepSeek V3.2DeepSeek146.383.4—164K$0.32
71Nemotron 3 UltraNVIDIA146.285.4—262K$1.05
72DeepSeek V4 Flash 0423DeepSeek146.1——1.05M$0.17
73Kimi K2 ThinkingMoonshot AI146.0——262K$1.07
74GLM 5Z.ai (Zhipu AI)145.887.8—205K$0.93
75MiniMax M2.7MiniMax145.8——205K$0.37
76GPT-5.4 NanoOpenAI145.878.5—400K$0.46
77o4 MiniOpenAI145.679.6—200K$1.93
78GPT-5 MiniOpenAI145.575.0—400K$0.69
79Gemini 2.5 Pro (Jun 2025)Google145.385.3———
80Gemini 3.5 Flash LiteGoogle145.183.3—1.05M$0.85
81DeepSeek V3.2 ExpDeepSeek145.0——164K$0.30
82Qwen3.7 FlashAlibaba (Qwen)144.682.3—1M$0.055
83Gemini 3.1 Flash LiteGoogle144.481.8—1.05M$0.56
84Grok 4 FastxAI144.2————
85Gemini 2.5 Pro (Mar 2025)Google144.2————
86Claude Opus 4.1Anthropic144.177.3—200K$30.00
87Qwen 3.5 Flash (hosted 35B-A3B)Alibaba (Qwen)144.0————
88Qwen 3.6 35B-A3BAlibaba (Qwen)143.9————
89Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba (Qwen)143.980.1———
90GLM 4.7Z.ai (Zhipu AI)143.583.3—205K$1.00
91Qwen 3.6 FlashAlibaba (Qwen)143.3————
92Gemini 2.5 Flash (Sep 2025)Google143.0————
93Gemma 4 31BGoogle142.775.8—262K$0.15
94Claude Opus 4Anthropic142.776.3———
95Qwen3.5-35B-A3BAlibaba (Qwen)142.583.5—262K$0.45
96GPT-5.5 InstantOpenAI142.582.5———
97Gemini 2.5 Pro (May 2025)Google142.5————
98Claude Haiku 4.5Anthropic142.471.2—200K$2.00
99Qwen3 MaxAlibaba (Qwen)142.4——262K$1.56
100o1OpenAI141.976.8—200K$26.25
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research