Skip to content
Benchmarks

11 benchmarks

All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
BenchmarkDomainResultsTop modelTop score
ARC AI2
ARC AI2 benchmark (score column: Challenge score).
other67DeepSeek V395.3%
Winogrande
Winogrande benchmark (score column: Accuracy).
other64Llama 3.1-405B89.2%
PIQA
PIQA benchmark (score column: Score).
other47GPT-4o-mini88.7%
Lech Mazur Writing
Lech Mazur Writing benchmark (score column: Mean score).
other46GPT-5 (medium)86.0%
TriviaQA
TriviaQA benchmark (score column: EM).
other33Llama 2-70B87.6%
OpenBookQA
OpenBookQA benchmark (score column: Accuracy).
other32phi-3-small 7.4B88.0%
ScienceQA
ScienceQA benchmark (score column: Score).
other22Phi-3.5-vision-instruct91.3%
LAMBADA
LAMBADA benchmark (score column: Score).
other21Falcon-180B79.8%
FrontierSWE
FrontierSWE benchmark (score column: Score).
other15GPT-6 Astra65.5%
ANLI
ANLI benchmark (score column: Score).
other9phi-3-small 7.4B58.1%
SuperGLUE
SuperGLUE benchmark (score column: Score).
other0—