Skip to content

Best for reasoning

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up.

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
101GPT-4.5 Preview (Feb 2025)OpenAI—68.7———
102DeepSeek-V3 (Mar 2025)DeepSeek135.967.6———
103Llama 4 Maverick (FP8)Meta—67.0———
104GPT-4.1OpenAI136.866.9—1.05M$3.50
105GPT-4.1 MiniOpenAI135.065.8—1.05M$0.70
106Qwen3 32BAlibaba (Qwen)138.565.7—131K$0.13
107Gemini 2.0 Pro Exp (Feb 2025)Google—65.7———
108QWQ-PlusAlibaba (Qwen)—65.4———
109QwQ-32BAlibaba (Qwen)137.665.3———
110DeepSeek-R1-Distill-Qwen-32BDeepSeek137.464.1———
111Gemini 2.0 Flash (Feb 2025)Google134.764.1———
112Qwen3 14BAlibaba (Qwen)138.263.8—131K$0.15
113o1-miniOpenAI · +1 variant135.862.4———
114Qwen3 30B A3BAlibaba (Qwen)136.261.7—131K$0.23
115gpt-oss-20bOpenAI137.860.8—131K$0.036
116GLM 4.7 FlashZ.ai (Zhipu AI)—60.5—200K$0.15
117Mistral Medium 3Mistral AI134.159.5—131K$0.80
118Gemini 1.5 Pro (Sept 2024)Google131.757.2———
119Gemini 2.0 Flash Thinking ExpGoogle—57.1———
120Qwen3 8BAlibaba (Qwen)136.256.8—131K$0.20
121DeepSeek V3DeepSeek132.456.5—164K$0.45
122Qwen2.5-MaxAlibaba (Qwen)132.556.1———
123Magistral Small 1.0Mistral AI133.256.1———
124Phi 4Microsoft130.456.1—16K$0.087
125DeepSeek-R1-Distill-Llama-70BDeepSeek—55.7———
126Qwen3-30B-A3B-Instruct (Jul 2025)Alibaba (Qwen)137.455.6———
127Claude 3.5 Sonnet (Oct 2024)Anthropic—55.3———
128Claude 3.5 Sonnet (Jun 2024)Anthropic—54.0———
129Grok-2 (Dec 2024)xAI130.553.8———
130Qwen3-4BAlibaba (Qwen)—52.3———
131Llama 4 ScoutMeta129.651.8—1.31M$0.15
132Mistral Large 2 (Nov 2024)Mistral AI128.551.3———
133Llama 3.1-405BMeta128.850.9———
134o1-previewOpenAI134.850.3———
135GPT-4o (Aug 2024)OpenAI128.849.2———
136Qwen2.5-72BAlibaba (Qwen)129.049.1———
137Mistral Small 3.2Mistral AI131.749.1———
138Mistral Large 2 (Jul 2024)Mistral AI127.549.0———
139GPT-4.1 NanoOpenAI129.648.9—1.05M$0.17
140GPT-4o (May 2024)OpenAI129.048.9———
141Qwen Plus (Jan 2025)Alibaba (Qwen)—48.1———
142GPT-4o (Nov 2024)OpenAI128.847.9———
143Gemma 3 27BGoogle130.047.7—131K$0.17
144Magistral Small 1.2Mistral AI131.447.6———
145Mistral Small 3.1Mistral AI127.547.5———
146Llama 3.3 70BMeta127.347.4———
147Gemini 1.5 Flash (Sep 2024)Google129.447.3———
148Mistral Small 3Mistral AI127.147.3—33K$0.058
149Claude 3 OpusAnthropic126.947.2———
150GPT-4 Turbo (Apr 2024)OpenAI127.346.6———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research