Skip to content

Newest models

Most recent releases, with whatever independent scores exist so far.

Most recent releases, with whatever independent scores exist so far. Source: OpenRouter / Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
201Llama 3.2 3BMeta—————
202Qwen2.5-72BAlibaba (Qwen)129.049.1———
203Qwen2.5-7BAlibaba (Qwen)118.435.5———
204Qwen2.5 72B InstructAlibaba (Qwen)———33K$0.37
205Qwen2.5-14BAlibaba (Qwen)—————
206Qwen2.5-Coder-32BAlibaba (Qwen)119.4————
207Qwen2.5-Coder (7B)Alibaba (Qwen)112.9————
208Qwen2.5-Coder (1.5B)Alibaba (Qwen)102.6————
209Qwen2.5-32BAlibaba (Qwen)128.546.1———
210Mistral Small v24.09Mistral AI—————
211Pixtral 12BMistral AI—————
212DeepSeek-V2.5 (Sep 2024)DeepSeek—————
213Command R+Cohere119.3————
214Llama 3.1 Euryale 70B v2.2Sao10K———131K$0.85
215Hermes 3 70B InstructNous Research———131K$0.70
216Phi-3.5-MoEMicrosoft—————
217Hermes 3 405B InstructNous Research———131K$1.00
218Phi-3.5-miniMicrosoft—————
219Llama 3 8B LunarisSao10K———8K$0.043
220Mistral Large 2 (Jul 2024)Mistral AI127.549.0———
221Llama 3.1-405BMeta128.850.9———
222Llama 3.1-70BMeta125.944.2———
223Llama 3.1-8BMeta116.527.0———
224Llama 3.1 70B InstructMeta———131K$0.40
225Llama 3.1 8B InstructMeta———131K$0.058
226Mistral NemoMistral AI118.629.9—131K$0.022
227Gemma 2 27BGoogle122.136.5—8K$0.65
228Gemma 2 9BGoogle119.827.5———
229Hermes 2 Theta Llama-3 70BNous Research—37.5———
230DeepSeek-Coder-V2 236BDeepSeek—————
231Qwen2-72BAlibaba (Qwen)125.340.8———
232Mistral 7B v0.3Mistral AI108.715.2———
233Yi-1.5-34B01.AI—32.0———
234Falcon 2 11BTechnology Innovation Institute109.3————
235DeepSeek-V2 (MoE-236B, May 2024)DeepSeek124.8————
236phi-3-small 7.4BMicrosoft121.8————
237phi-3-medium 14BMicrosoft121.227.6———
238phi-3-mini 3.8BMicrosoft117.2————
239Llama 3-70BMeta122.940.6———
240Llama 3-8BMeta116.326.1———
241Mixtral 8x22BMistral AI122.034.1———
242Mixtral 8x22B InstructMistral AI———66K$3.00
243WizardLM-2 8x22BMicrosoft—43.4—66K$0.62
244Qwen1.5-32BAlibaba (Qwen)—30.7———
245DBRXDatabricks—32.9———
246StarCoder 2 3BHugging Face88.0————
247Gemma 7BGoogle111.7————
248Gemma 2BGoogle93.6————
249StarCoder 2 15BHugging Face104.6————
250StarCoder 2 7BHugging Face92.9————
— means no published score from that source yet · click a model for every benchmark with its source
Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research