Skip to content
Benchmarks

3 benchmarks

All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
BenchmarkDomainResultsTop modelTop score
MMLU
MMLU benchmark (score column: EM).
knowledge119GPT-4o (Nov 2024)88.1%
SimpleQA Verified
Short factual questions testing parametric knowledge and hallucination.
knowledge78GPT-6 Astra75.6%
HLE
Humanity's Last Exam: expert-written questions across many subjects.
knowledge48GPT-6 Astra (unknown thinking)54.8%