| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Claude Mythos Preview (Early) Anthropic | 73.8% | — | 7 Apr 2026 | — | Source ↗ | |
| 2 | GPT-5.5 (unknown thinking) OpenAI | 47.4% | unknown | 23 Apr 2026 | — | Source ↗ | |
| 3 | Claude Opus 4.7 (no thinking) Anthropic | 26.5% | — | 16 Apr 2026 | — | Source ↗ | |
| 4 | Gemini 3.1 Pro Preview Google | 26.1% | — | 19 Feb 2026 | — | Source ↗ | |
| 5 | Claude Sonnet 4.6 (no thinking) Anthropic | 23.6% | — | 17 Feb 2026 | — | Source ↗ | |
| 6 | Kimi K2.6 Open source MoonshotAI | 18.4% | — | 20 Apr 2026 | — | Source ↗ | |
| 7 | GLM 5.1 Open source Z.ai | 18.1% | — | 7 Apr 2026 | — | Source ↗ | |
| 8 | Claude Haiku 4.5 (unknown thinking) Anthropic | 13.7% | unknown | 15 Oct 2025 | — | Source ↗ | |
| 9 | MiniMax M2.7 Open source MiniMax | 13.3% | — | 18 Mar 2026 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.