Skip to content

Best for coding

Ranked by DeepSWE - agentic software-engineering tasks on real repositories.

Ranked by DeepSWE - agentic software-engineering tasks on real repositories. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
1GPT-6 AstraNEWxhighOpenAI · +3 variants——74.1——
2Gemini 3.8 FlashNEWGoogle · +1 variant157.195.473.81.05M$1.50
3Claude Opus 5Anthropic · +3 variants162.793.973.61M$10.00
4GPT-5.6 SolOpenAI · +3 variants162.093.572.71.05M$4.00
5GPT-5.6 TerraOpenAI · +3 variants159.393.369.61.05M$4.50
6GLM 5.3Z.ai (Zhipu AI)155.690.969.01.31M$2.15
7Kimi K3Moonshot AI157.793.168.51.05M$6.00
8Grok 4.6mediumxAI · +3 variants——67.5——
9GPT-5.6 LunaOpenAI · +3 variants156.391.667.21.05M$0.45
10Gemini 3.7 FlashmediumGoogle · +2 variants——65.5——
11GLM 5.3 FlashZ.ai (Zhipu AI)151.990.263.41.31M$0.24
12Qwen3.8 MaxxhighAlibaba (Qwen)—92.757.5——
13Muse Spark 1.2xhighMeta——54.9——
14Grok 4.5xAI154.093.453.8500K$3.00
15Muse Spark 1.1Meta154.3—53.31.05M$2.00
16Gemini 3.6 FlashGoogle154.494.146.71.05M$1.50
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research