Benchmarks
6 benchmarks
All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
| Benchmark | Domain | Results | Top model | Top score | Released | Latest run | Type |
|---|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 Competition math problems from OTIS mock AIME exams. | math | 244 | GPT-6 Astra | 100.0% | 19 Dec 2024 | 17 Sep 2026 | Run by Epoch |
| MATH level 5 The hardest subset of the MATH competition dataset. | math | 104 | GPT-5 | 98.1% | 5 Mar 2021 | 30 Oct 2025 | Run by Epoch |
| FrontierMath-Tiers-1-3-v2-Private Unpublished, expert-written research-level mathematics problems (tiers 1-3). | math | 95 | GPT-6 Astra | 93.7% | 12 Jun 2026 | 18 Sep 2026 | Run by Epoch |
| ProofBench ProofBench benchmark (score column: Accuracy). | math | 70 | Claude Sonnet 5.5 | 100.0% | 30 Jan 2026 | — | External leaderboard |
| GSM8K GSM8K benchmark (score column: EM). | math | 68 | DeepSeek-Coder-V2 236B | 94.5% | 27 Oct 2021 | — | External leaderboard |
| FrontierMath-Tier-4-v2-Private The hardest FrontierMath tier: research problems that take experts days. | math | 63 | GPT-6 Astra | 97.6% | 12 Jun 2026 | 18 Sep 2026 | Run by Epoch |