Skip to content

Best for math

Ranked by Epoch's mock AIME 2024-2025 (competition mathematics).

Ranked by Epoch's mock AIME 2024-2025 (competition mathematics). Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
151Mistral Small 3Mistral AI127.147.3—33K$0.058
152Claude 3.5 Sonnet (Jun 2024)Anthropic—54.0———
153GPT-4o (Aug 2024)OpenAI128.849.2———
154GPT-4o (May 2024)OpenAI129.048.9———
155GPT-4o (Nov 2024)OpenAI128.847.9———
156Qwen-TurboAlibaba (Qwen)—41.8———
157Mistral Small 3.1Mistral AI127.547.5———
158Llama 3.3 70BMeta127.347.4———
159Claude 3 OpusAnthropic126.947.2———
160Gemini 1.5 Flash 8BGoogle—33.0———
161Tulu 3 (Tülu 3) 70BAllen Institute for AI—46.3———
162Llama 3-70BMeta122.940.6———
163Claude 3.5 Haiku (Oct 2024)Anthropic—38.1———
164Gemini 1.5 Flash (May 2024)Google122.640.4———
165Llama 3.1-70BMeta125.944.2———
166Granite 4.0 MicroIBM—28.3—131K$0.041
167Llama 3.2 90BMeta125.541.0———
168Claude 3 SonnetAnthropic120.740.6———
169Claude 2Anthropic120.134.7———
170Qwen2.5-7BAlibaba (Qwen)118.435.5———
171Hermes 2 Theta Llama-3 70BNous Research—37.5———
172GPT-3.5 Turbo (Jan 2024)OpenAI115.627.2———
173Mistral LargeMistral AI122.038.8—128K$3.00
174Claude 2.1Anthropic119.233.0———
175Llama 3-8BMeta116.326.1———
176Claude 3 HaikuAnthropic118.336.3———
177Llama 3.1-8BMeta116.527.0———
178Gemma 2 27BGoogle122.136.5—8K$0.65
179GPT-4 (Jun 2023)OpenAI123.130.7———
180Gemini 1.0 ProGoogle117.034.0———
181Gemma 3 1BGoogle—19.9———
182DeepSeek LLM 67BDeepSeek110.524.6———
183granite-4.0-1b—24.0———
184GPT-4 (Mar 2023)OpenAI125.935.7———
185Gemma 2 9BGoogle119.827.5———
186Llama 3.2 1BMeta102.023.9———
187Mistral 7B v0.3Mistral AI108.715.2———
188granite-4.0-350m—11.2———
189Llama 2-70BMeta113.626.3———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research