| Benchmark | Domain | Results | Top model | Top score | Released | Latest run | Type |
|---|---|---|---|---|---|---|---|
| WeirdML Novel, unusual machine-learning tasks solved by writing and running code. | coding | 149 | GPT-6 Astra (pro, max) | 93.6% | 16 Jan 2025 | — | External leaderboard |
| Aider polyglot Code editing exercises in C++, Go, Java, JavaScript, Python and Rust. | coding | 60 | GPT-5 | 88.0% | 21 Dec 2024 | — | External leaderboard |
| DeepSWE DeepSWE benchmark (score column: Pass@1). | coding | 56 | GPT-6 Astra (xhigh) | 74.1% | 26 May 2026 | — | External leaderboard |
| GSO-Bench GSO-Bench benchmark (score column: Score OPT@1). | coding | 36 | Claude Opus 4.8 | 47.1% | 29 May 2025 | — | External leaderboard |
| FrontierCode FrontierCode benchmark (score column: Main score). | coding | 36 | Claude Fable 5 (unknown) | 53.5% | 8 Jun 2026 | — | External leaderboard |
| SWE-Bench verified 500 human-validated real GitHub issues; the model must produce a patch that passes tests. | coding | 33 | Claude Opus 4.7 | 83.5% | 13 Aug 2024 | 25 Jun 2026 | Run by Epoch |
| Cybench Cybench benchmark (score column: Unguided % Solved). | coding | 22 | Claude Opus 4.6 (unknown thinking) | 93.0% | 15 Aug 2024 | — | External leaderboard |
| CadEval CadEval benchmark (score column: Overall pass (%)). | coding | 14 | o3 (medium) | 74.0% | 22 Apr 2025 | — | External leaderboard |
| ExploitBench ExploitBench benchmark (score column: Mean capability). | coding | 9 | Claude Mythos Preview (Early) | 73.8% | 13 May 2026 | — | External leaderboard |
| MirrorCode MirrorCode benchmark (score column: Best score (across scorers)). | coding | 8 | Claude Fable 5.1 | 73.3% | 26 Jun 2026 | 10 Sep 2026 | Run by Epoch |