Updated 2h ago
AI Leaderboard - 916 models ranked by capability, price & context
Independent benchmark results from Epoch AI and Artificial Analysis, live API prices and context windows from OpenRouter. Each ranking uses one named measure - we don't blend them into a made-up score.
Ranked by the Epoch Capabilities Index (ECI) - Epoch AI's statistical model over dozens of benchmarks. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.
| # | Model | Capability | Reasoning | Coding | Context | $/M | Weights | Released |
|---|---|---|---|---|---|---|---|---|
| 1 | Kimi K3Moonshot AI | 157.7 | 93.1 | 68.5 | 1.05M | $6.00 | Open | 16 Jul 2026 |
| 2 | GLM 5.3Z.ai (Zhipu AI) | 155.6 | 90.9 | 69.0 | 1.31M | $2.15 | Open | 14 Aug 2026 |
| 3 | DeepSeek V4 Pro 0813DeepSeek | 155.4 | 91.7 | — | 1.05M | $1.41 | Open | 13 Aug 2026 |
| 4 | DeepSeek V4.1 FlashNEWDeepSeek | 155.0 | — | — | 1.05M | $0.53 | Open | 9 Sep 2026 |
| 5 | DeepSeek V4 Flash 0731DeepSeek | 154.5 | 91.0 | — | 1.31M | $0.093 | Open | 31 Jul 2026 |
| 6 | GLM 5.3 FlashZ.ai (Zhipu AI) | 151.9 | 90.2 | 63.4 | 1.31M | $0.24 | Open | 20 Aug 2026 |
| 7 | Inkling SmallThinking Machines Lab | 150.1 | — | — | 524K | $0.64 | Open | 15 Jul 2026 |
| 8 | Qwen 3.8 27BAlibaba (Qwen) | 149.4 | — | — | — | — | Open | 14 Aug 2026 |
| 9 | InklingThinking Machines Lab | 148.6 | — | — | 524K | $1.76 | Open | 15 Jul 2026 |
— means no published score from that source yet · click a model for every benchmark with its source
New - awaiting independent evaluation
All new models →- Claude Sonnet 5.5Anthropic · 28 Sep 2026 · AA 56 · 1M · $4.00/M
- Qwen3.8 Max PrimeAlibaba (Qwen) · 23 Sep 2026 · 1M · $6.00/M
- GPT-6 SolOpenAI · 22 Sep 2026 · AA 47.5 · 1.05M · $4.00/M
- GPT-6 LunaOpenAI · 22 Sep 2026 · AA 37.3 · 1.05M · $0.20/M
- GPT-6 Sol ProOpenAI · 22 Sep 2026 · 1.05M · $4.00/M
- GPT-6 Luna ProOpenAI · 22 Sep 2026 · 1.05M · $0.20/M
- Claude Opus 5.5Anthropic · 22 Sep 2026 · AA 57.6 · 1M · $8.00/M
- Grok 4.7xAI · 21 Sep 2026 · AA 46.4 · 500K · $3.00/M
Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.
Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research