Skip to content
Benchmarks

6 benchmarks

All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
BenchmarkDomainResultsTop modelTop score
OTIS Mock AIME 2024-2025
Competition math problems from OTIS mock AIME exams.
math244GPT-6 Astra100.0%
MATH level 5
The hardest subset of the MATH competition dataset.
math104GPT-598.1%
FrontierMath-Tiers-1-3-v2-Private
Unpublished, expert-written research-level mathematics problems (tiers 1-3).
math95GPT-6 Astra93.7%
ProofBench
ProofBench benchmark (score column: Accuracy).
math70Claude Sonnet 5.5100.0%
GSM8K
GSM8K benchmark (score column: EM).
math68DeepSeek-Coder-V2 236B94.5%
FrontierMath-Tier-4-v2-Private
The hardest FrontierMath tier: research problems that take experts days.
math63GPT-6 Astra97.6%