Skip to content
Benchmarks

10 benchmarks

All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
BenchmarkDomainResultsTop modelTop score
WeirdML
Novel, unusual machine-learning tasks solved by writing and running code.
coding149GPT-6 Astra (pro, max)93.6%
Aider polyglot
Code editing exercises in C++, Go, Java, JavaScript, Python and Rust.
coding60GPT-588.0%
DeepSWE
DeepSWE benchmark (score column: Pass@1).
coding56GPT-6 Astra (xhigh)74.1%
FrontierCode
FrontierCode benchmark (score column: Main score).
coding39Claude Fable 5 (unknown)53.5%
GSO-Bench
GSO-Bench benchmark (score column: Score OPT@1).
coding36Claude Opus 4.847.1%
SWE-Bench verified
500 human-validated real GitHub issues; the model must produce a patch that passes tests.
coding33Claude Opus 4.783.5%
Cybench
Cybench benchmark (score column: Unguided % Solved).
coding22Claude Opus 4.6 (unknown thinking)93.0%
CadEval
CadEval benchmark (score column: Overall pass (%)).
coding14o3 (medium)74.0%
ExploitBench
ExploitBench benchmark (score column: Mean capability).
coding9Claude Mythos Preview (Early)73.8%
MirrorCode
MirrorCode benchmark (score column: Best score (across scorers)).
coding8Claude Fable 5.173.3%