Skip to content

Best for math

Ranked by Epoch's mock AIME 2024-2025 (competition mathematics).

Ranked by Epoch's mock AIME 2024-2025 (competition mathematics). Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
101Qwen3-30B-A3B-Instruct (Jul 2025)Alibaba (Qwen)137.455.6———
102Qwen3.5-9BAlibaba (Qwen)139.479.0—262K$0.11
103QwQ-32BAlibaba (Qwen)137.665.3———
104GLM 4.7 FlashZ.ai (Zhipu AI)—60.5—200K$0.15
105Claude 3.7 Sonnet64k thinkingAnthropic · +3 variants—79.7———
106Gemini 2.0 Flash Thinking ExpGoogle—57.1———
107Qwen3 8BAlibaba (Qwen)136.256.8—131K$0.20
108Qwen3.5-4BAlibaba (Qwen)—————
109Grok 3xAI138.375.8———
110DeepSeek-R1-Distill-Qwen-32BDeepSeek137.464.1———
111R1DeepSeek139.071.7—64K$1.15
112qwen3-4b-instruct-2507—45.8———
113DeepSeek-R1-Distill-Llama-70BDeepSeek—55.7———
114DeepSeek-R1-Distill-Qwen-14BDeepSeek135.444.7———
115o1-miniOpenAI · +1 variant135.862.4———
116GPT-4.1 MiniOpenAI135.065.8—1.05M$0.70
117deepseek-r1-0528-qwen3-8b—9.3———
118GPT-4.1OpenAI136.866.9—1.05M$3.50
119DeepSeek-V3 (Mar 2025)DeepSeek135.967.6———
120GPT-4.5 Preview (Feb 2025)OpenAI—68.7———
121Mistral Medium 3Mistral AI134.159.5—131K$0.80
122o1-previewOpenAI134.850.3———
123Gemini 2.0 Flash (Feb 2025)Google134.764.1———
124Mistral Small 3.2Mistral AI131.749.1———
125Magistral Small 1.0Mistral AI133.256.1———
126GPT-4.1 NanoOpenAI129.648.9—1.05M$0.17
127Magistral Small 1.2Mistral AI131.447.6———
128Gemini 1.5 Pro (Sept 2024)Google131.757.2———
129Gemma 3 27BGoogle130.047.7—131K$0.17
130DeepSeek-R1-Distill-Qwen-1.5BDeepSeek—33.6———
131Llama 4 Maverick (FP8)Meta—67.0———
132Qwen Plus (Jan 2025)Alibaba (Qwen)—48.1———
133Gemma 3 12BGoogle123.539.5—131K$0.075
134Gemini 1.5 Flash (Sep 2024)Google129.447.3———
135Qwen2.5-MaxAlibaba (Qwen)132.556.1———
136DeepSeek V3DeepSeek132.456.5—164K$0.45
137Phi 4Microsoft130.456.1—16K$0.087
138Grok-2 (Dec 2024)xAI130.553.8———
139Llama 3.1-405BMeta128.850.9———
140Mistral Large 2 (Jul 2024)Mistral AI127.549.0———
141Claude 3.5 Sonnet (Oct 2024)Anthropic—55.3———
142Qwen2.5-72BAlibaba (Qwen)129.049.1———
143Qwen3-1.7BAlibaba (Qwen)—38.0———
144Llama 4 ScoutMeta129.651.8—1.31M$0.15
145Mistral Large 2 (Nov 2024)Mistral AI128.551.3———
146Gemma 3 4BGoogle116.023.2—131K$0.063
147Qwen2.5-32BAlibaba (Qwen)128.546.1———
148GPT-4o-miniOpenAI126.637.7—128K$0.26
149Gemini 1.5 Pro (May 2024)Google126.945.9———
150GPT-4 Turbo (Apr 2024)OpenAI127.346.6———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research