Benchmarks
11 benchmarks
All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
| Benchmark | Domain | Results | Top model | Top score | Released | Latest run | Type |
|---|---|---|---|---|---|---|---|
| ARC AI2 ARC AI2 benchmark (score column: Challenge score). | other | 67 | DeepSeek V3 | 95.3% | 14 Mar 2018 | — | External leaderboard |
| Winogrande Winogrande benchmark (score column: Accuracy). | other | 64 | Llama 3.1-405B | 89.2% | 24 Jul 2019 | — | External leaderboard |
| PIQA PIQA benchmark (score column: Score). | other | 47 | GPT-4o-mini | 88.7% | 26 Nov 2019 | — | External leaderboard |
| Lech Mazur Writing Lech Mazur Writing benchmark (score column: Mean score). | other | 46 | GPT-5 (medium) | 86.0% | 31 Jan 2025 | — | External leaderboard |
| TriviaQA TriviaQA benchmark (score column: EM). | other | 33 | Llama 2-70B | 87.6% | 9 May 2017 | — | External leaderboard |
| OpenBookQA OpenBookQA benchmark (score column: Accuracy). | other | 32 | phi-3-small 7.4B | 88.0% | 8 Sep 2018 | — | External leaderboard |
| ScienceQA ScienceQA benchmark (score column: Score). | other | 22 | Phi-3.5-vision-instruct | 91.3% | 20 Sep 2022 | — | External leaderboard |
| LAMBADA LAMBADA benchmark (score column: Score). | other | 21 | Falcon-180B | 79.8% | 20 Jun 2016 | — | External leaderboard |
| FrontierSWE FrontierSWE benchmark (score column: Score). | other | 15 | GPT-6 Astra | 65.5% | 2 Sep 2026 | — | External leaderboard |
| ANLI ANLI benchmark (score column: Score). | other | 9 | phi-3-small 7.4B | 58.1% | 31 Oct 2019 | — | External leaderboard |
| SuperGLUE SuperGLUE benchmark (score column: Score). | other | 0 | — | 2 May 2019 | — | External leaderboard |