Skip to content
Benchmarks

58 benchmarks

All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
BenchmarkDomainResultsTop modelTop score
GPQA diamond
198 graduate-level, 'Google-proof' multiple-choice questions in biology, physics and chemistry.
science262GPT-6 Astra95.8%
OTIS Mock AIME 2024-2025
Competition math problems from OTIS mock AIME exams.
math244GPT-6 Astra100.0%
DTBench
DTBench benchmark (score column: Accuracy).
reasoning205Claude Opus 5.598.9%
Chess Puzzles
Chess tactics puzzles solved by the model without tools.
games188GPT-6 Astra72.0%
ARC-AGI
ARC-AGI benchmark (score column: Score).
reasoning176GPT-6 Astra98.5%
ARC-AGI-2
Abstract visual reasoning puzzles designed to be easy for humans and hard for AI.
reasoning170GPT-6 Astra95.0%
LMCA
LMCA benchmark (score column: Score).
agents167Claude Opus 5.568.2%
WeirdML
Novel, unusual machine-learning tasks solved by writing and running code.
coding149GPT-6 Astra (pro, max)93.6%
MMLU
MMLU benchmark (score column: EM).
knowledge119GPT-4o (Nov 2024)88.1%
Mystery Game Puzzles
Mystery Game Puzzles benchmark (score column: Best score (across scorers)).
games107GPT-6 Astra84.0%
MATH level 5
The hardest subset of the MATH competition dataset.
math104GPT-598.1%
SimpleBench
Trick questions on spatio-temporal and social reasoning where humans outperform models.
reasoning98Claude Fable 581.9%
FrontierMath-Tiers-1-3-v2-Private
Unpublished, expert-written research-level mathematics problems (tiers 1-3).
math95GPT-6 Astra93.7%
SimpleQA Verified
Short factual questions testing parametric knowledge and hallucination.
knowledge78GPT-6 Astra75.6%
ProofBench
ProofBench benchmark (score column: Accuracy).
math70Claude Sonnet 5.5100.0%
GSM8K
GSM8K benchmark (score column: EM).
math68DeepSeek-Coder-V2 236B94.5%
ARC AI2
ARC AI2 benchmark (score column: Challenge score).
other67DeepSeek V395.3%
Winogrande
Winogrande benchmark (score column: Accuracy).
other64Llama 3.1-405B89.2%
FrontierMath-Tier-4-v2-Private
The hardest FrontierMath tier: research problems that take experts days.
math63GPT-6 Astra97.6%
Aider polyglot
Code editing exercises in C++, Go, Java, JavaScript, Python and Rust.
coding60GPT-588.0%
Terminal Bench
Agentic tasks completed in a real terminal environment (external leaderboard).
agents58GPT-5.5 (unknown thinking)84.7%
Fiction.LiveBench
Long-context comprehension of fiction stories.
long-context57GPT-5 (medium)97.2%
DeepSWE
DeepSWE benchmark (score column: Pass@1).
coding56GPT-6 Astra (xhigh)74.1%
HellaSwag
HellaSwag benchmark (score column: Overall accuracy).
reasoning52GPT-4 (Mar 2023)95.3%
HLE
Humanity's Last Exam: expert-written questions across many subjects.
knowledge48GPT-6 Astra (unknown thinking)54.8%
PIQA
PIQA benchmark (score column: Score).
other47GPT-4o-mini88.7%
Lech Mazur Writing
Lech Mazur Writing benchmark (score column: Mean score).
other46GPT-5 (medium)86.0%
METR Time Horizons
Length of software tasks (in human time) an AI agent completes with 50% reliability.
agents46Claude Mythos Preview (Early)85.2%
BBH
BBH benchmark (score column: Average).
reasoning45Gemini 1.5 Pro (May 2024)89.2%
Balrog
Balrog benchmark (score column: Average progress).
games39GPT-6 Astra68.3%
FrontierCode
FrontierCode benchmark (score column: Main score).
coding39Claude Fable 5 (unknown)53.5%
APEX-Agents
APEX-Agents benchmark (score column: Pass@1 score).
agents38Claude Sonnet 5.575.5%
GSO-Bench
GSO-Bench benchmark (score column: Score OPT@1).
coding36Claude Opus 4.847.1%
DeepResearch Bench
DeepResearch Bench benchmark (score column: Average score).
agents35Claude Opus 4.655.3%
VPCT
VPCT benchmark (score column: Correct).
multimodal34Gemini 3 Pro Preview91.0%
SWE-Bench verified
500 human-validated real GitHub issues; the model must produce a patch that passes tests.
coding33Claude Opus 4.783.5%
TriviaQA
TriviaQA benchmark (score column: EM).
other33Llama 2-70B87.6%
OpenBookQA
OpenBookQA benchmark (score column: Accuracy).
other32phi-3-small 7.4B88.0%
GeoBench
GeoBench benchmark (score column: ACW Country %).
multimodal30Gemini 3 Flash Preview88.0%
Surface Evolver Bench
Surface Evolver Bench benchmark (score column: Mean score).
science27Kimi K3 (unknown)95.0%
Furniture Assembly
Furniture Assembly benchmark (score column: Best score (across scorers)).
multimodal27Claude Opus 5.583.3%
EBR-bench
EBR-bench benchmark (score column: Best score (across scorers)).
reasoning23GPT-6 Astra76.2%
Cybench
Cybench benchmark (score column: Unguided % Solved).
coding22Claude Opus 4.6 (unknown thinking)93.0%
ScienceQA
ScienceQA benchmark (score column: Score).
other22Phi-3.5-vision-instruct91.3%
CL-bench
CL-bench benchmark (score column: Overall).
long-context22GPT-5.4 (xhigh)27.9%
LAMBADA
LAMBADA benchmark (score column: Score).
other21Falcon-180B79.8%
CL-bench Life
CL-bench Life benchmark (score column: Overall).
long-context16GPT-5.522.2%
Remote Labor Index
Remote Labor Index benchmark (score column: Score).
agents15GPT-6 Astra (unknown thinking)20.8%
CadEval
CadEval benchmark (score column: Overall pass (%)).
coding14o3 (medium)74.0%
The Agent Company
The Agent Company benchmark (score column: % Resolved).
agents14DeepSeek V3.2 Exp42.9%
OSWorld 2.0
OSWorld 2.0 benchmark (score column: Binary accuracy).
agents13Claude Opus 531.4%
PostTrainBench
PostTrainBench benchmark (score column: Average (%)).
agents11Claude Fable 541.8%
GDPval
Economically valuable tasks across occupations, judged against expert work.
agents11GPT-5.2 (none)49.7%
ANLI
ANLI benchmark (score column: Score).
other9phi-3-small 7.4B58.1%
OSWorld
Computer-use tasks in real desktop operating systems.
agents9Claude Sonnet 4.6 (no thinking)72.1%
ExploitBench
ExploitBench benchmark (score column: Mean capability).
coding9Claude Mythos Preview (Early)73.8%
MirrorCode
MirrorCode benchmark (score column: Best score (across scorers)).
coding8Claude Fable 5.173.3%
SuperGLUE
SuperGLUE benchmark (score column: Score).
other0—