Skip to content

Open-weight leaderboard

Models whose weights are published, ranked by the Epoch Capabilities Index.

Models whose weights are published, ranked by the Epoch Capabilities Index. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
101Phi-2Microsoft107.6————
102LLaMA-33BMeta107.1————
103Qwen-7BAlibaba (Qwen)106.5————
104Llama 2-13BMeta105.8————
105StarCoder 2 15BHugging Face104.6————
106Yi 6B01.AI104.4————
107Falcon-40BTechnology Innovation Institute104.0————
108Baichuan2-13BBaichuan102.8————
109Qwen2.5-Coder (1.5B)Alibaba (Qwen)102.6————
110Llama 3.2 1BMeta102.023.9———
111MPT-30BMosaicML100.2————
112LLaMA-13BMeta100.1————
113Llama 2-7BMeta98.5————
114LLaMA-7BMeta96.1————
115Baichuan 2-7BBaichuan95.8————
116DeepSeek Coder 33BDeepSeek95.7————
117Falcon-7BTechnology Innovation Institute94.5————
118MPT-7BMosaicML94.0————
119Gemma 2BGoogle93.6————
120StarCoder 2 7BHugging Face92.9————
121XGen-7BSalesforce92.8————
122Phi-1.5Microsoft90.8————
123Baichuan1-7BBaichuan89.8————
124Dolly 2.0-12bDatabricks88.9————
125DeepSeek Coder 6.7BDeepSeek88.9————
126StarCoder 2 3BHugging Face88.0————
127Cerebras-GPT-13BCerebras82.4————
128DeepSeek Coder 1.3BDeepSeek62.2————
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research