Skip to content

Best for reasoning

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up.

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
51Llama 3.1-405BMeta128.850.9———
52Qwen2.5-72BAlibaba (Qwen)129.049.1———
53Mistral Small 3.2Mistral AI131.749.1———
54Mistral Large 2 (Jul 2024)Mistral AI127.549.0———
55Gemma 3 27BGoogle130.047.7—131K$0.17
56Magistral Small 1.2Mistral AI131.447.6———
57Mistral Small 3.1Mistral AI127.547.5———
58Llama 3.3 70BMeta127.347.4———
59Mistral Small 3Mistral AI127.147.3—33K$0.058
60Tulu 3 (Tülu 3) 70BAllen Institute for AI—46.3———
61Qwen2.5-32BAlibaba (Qwen)128.546.1———
62DeepSeek-R1-Distill-Qwen-14BDeepSeek135.444.7———
63Llama 3.1-70BMeta125.944.2———
64WizardLM-2 8x22BMicrosoft—43.4—66K$0.62
65Llama 3.2 90BMeta125.541.0———
66Qwen2-72BAlibaba (Qwen)125.340.8———
67Llama 3-70BMeta122.940.6———
68Gemma 3 12BGoogle123.539.5—131K$0.075
69Qwen3-1.7BAlibaba (Qwen)—38.0———
70Hermes 2 Theta Llama-3 70BNous Research—37.5———
71Gemma 2 27BGoogle122.136.5—8K$0.65
72Qwen2.5-7BAlibaba (Qwen)118.435.5———
73Mixtral 8x22BMistral AI122.034.1———
74Eurus-2-7B-PRIMETsinghua University,University of Illinois Urbana-Champaign (UIUC),Shanghai AI Lab,Peking University,Shanghai Jiao Tong University,CUHK Shenzhen Research Institute—33.9———
75DeepSeek-R1-Distill-Qwen-1.5BDeepSeek—33.6———
76DBRXDatabricks—32.9———
77Yi-1.5-34B01.AI—32.0———
78Qwen1.5-32BAlibaba (Qwen)—30.7———
79Mixtral 8x7BMistral AI118.430.6———
80Mistral NemoMistral AI118.629.9—131K$0.022
81Qwen1.5-72BAlibaba (Qwen)—28.8———
82Granite 4.0 MicroIBM—28.3—131K$0.041
83phi-3-medium 14BMicrosoft121.227.6———
84Gemma 2 9BGoogle119.827.5———
85Ministral 8BMistral AI—27.1———
86Llama 3.1-8BMeta116.527.0———
87Llama 2-70BMeta113.626.3———
88Llama 3-8BMeta116.326.1———
89DeepSeek LLM 67BDeepSeek110.524.6———
90Llama 3.2 1BMeta102.023.9———
91Gemma 3 4BGoogle116.023.2—131K$0.063
92Gemma 3 1BGoogle—19.9———
93Mistral 7B v0.3Mistral AI108.715.2———
94Yi-34B01.AI117.314.7———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research