Newest models
Most recent releases, with whatever independent scores exist so far.
Most recent releases, with whatever independent scores exist so far. Source: OpenRouter / Epoch AI. Variants (reasoning effort, batch) are folded into one row.
| # | Model | Capability | Reasoning | Coding | Context | $/M | Weights | Released |
|---|---|---|---|---|---|---|---|---|
| 651 | GPT-4 (Mar 2023)OpenAI | 125.9 | 35.7 | — | — | — | Closed | 14 Mar 2023 |
| 652 | LLaMA-65BMeta | 109.9 | — | — | — | — | Open | 24 Feb 2023 |
| 653 | LLaMA-33BMeta | 107.1 | — | — | — | — | Open | 24 Feb 2023 |
| 654 | LLaMA-13BMeta | 100.1 | — | — | — | — | Open | 24 Feb 2023 |
| 655 | LLaMA-7BMeta | 96.1 | — | — | — | — | Open | 24 Feb 2023 |
| 656 | BLIP-2 (Q-Former)Salesforce Research | — | — | — | — | — | Open | 6 Feb 2023 |
— means no published score from that source yet · click a model for every benchmark with its source
Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research