Skip to content

Newest models

Most recent releases, with whatever independent scores exist so far.

Most recent releases, with whatever independent scores exist so far. Source: OpenRouter / Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
551Mixtral 8x22B InstructMistral AI———66K$3.00
552WizardLM-2 8x22BMicrosoft—43.4—66K$0.62
553CodeQwen1.5-7B94.2————
554GPT-4 Turbo (Apr 2024)OpenAI127.346.6———
555GPT-4 TurboOpenAI · +1 variant———128K$15.00
556Qwen1.5-32BAlibaba (Qwen)—30.7———
557DBRXDatabricks—32.9———
558openhands-lm-32b-v0.1—————
559MM1-3B-Chat—————
560MM1-7B-Chat—————
561Claude 3 HaikuAnthropic118.336.3———
562Yi-9B107.3————
563Claude 3 OpusAnthropic126.947.2———
564Claude 3 SonnetAnthropic120.740.6———
565Mistral LargeMistral AI122.038.8—128K$3.00
566Nemotron-4 15BNVIDIA107.4————
567mistral-small-2402—————
568StarCoder 2 3BHugging Face88.0————
569Gemma 7BGoogle111.7————
570Gemma 2BGoogle93.6————
571StarCoder 2 15BHugging Face104.6————
572StarCoder 2 7BHugging Face92.9————
573Gemini 1.5 Pro (Feb 2024)Google—————
574Qwen1.5-14BAlibaba (Qwen)—————
575Qwen1.5-72BAlibaba (Qwen)—28.8———
576Qwen1.5-7BAlibaba (Qwen)—————
577llava-v1.6-mistral-7b—————
578llava-v1.6-vicuna-13b—————
579llava-v1.6-vicuna-7b—————
580GPT-4 Turbo (Nov 2023)OpenAI126.5————
581GPT-3.5 Turbo (Jan 2024)OpenAI115.627.2———
582GPT-3.5 Turbo (older v0613)OpenAI———4K$1.25
583GPT-4 Turbo Preview (January 2024)OpenAI—42.3———
584Gemini 1.0 Pro VisionGoogle—————
585instructblip-vicuna-13b—————
586InternVL-Chat-ViT-6B-Vicuna-7B—————
587Gemini 1.0 ProGoogle117.034.0———
588Phi-2Microsoft107.6————
589Mixtral 8x7BMistral AI118.430.6———
590Mistral 7B v0.2Mistral AI—————
591Mistral MediumMistral AI—————
592Qwen-1_8B92.3————
593DeepSeek LLM 67BDeepSeek110.524.6———
594Yi 6B01.AI104.4————
595Claude 2.1Anthropic119.233.0———
596GPT-3.5 Turbo (Nov 2023)OpenAI118.528.0———
597GPT-4 Turbo Preview (Nov 2023)OpenAI—42.4———
598Yi-34B01.AI117.314.7———
599DeepSeek Coder 33BDeepSeek95.7————
600DeepSeek Coder 6.7BDeepSeek88.9————
— means no published score from that source yet · click a model for every benchmark with its source
Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research