Skip to content

Best for agents

Ranked by APEX-Agents - long-horizon professional tasks performed by agents.

Ranked by APEX-Agents - long-horizon professional tasks performed by agents. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
1GLM-5.3-FlashunknownZ.ai (Zhipu AI)—————
2Kimi K3unknownMoonshot AI—————
3DeepSeek V4 Pro 0813unknownDeepSeek—————
4GLM 5.1Z.ai (Zhipu AI)149.989.9—205K$2.15
5MiniMax M3MiniMax147.090.9—1.05M$0.53
6Kimi K2.7 CodeMoonshot AI150.087.930.5262K$1.32
7InklingThinking Machines Lab148.6——524K$1.76
8Qwen3.5 397B A17BAlibaba (Qwen)146.686.4—262K$1.29
9gpt-oss-120bOpenAI139.975.8—131K$0.070
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research